GLM 5.3 Flash
Description
GLM-5.3-Flash is a native multimodal model from Z.ai, optimized for efficient coding and long-horizon agent tasks. Built on a 320B-parameter MoE architecture with 18B active parameters, it supports text, image, and video input with a hybrid sparse and linear attention design that maintains accurate long-context behavior while reducing compute overhead. Despite its agent-first positioning, it is surprisingly good at roleplay.
At a Glance
Key pricing and model details available for this model.
Input price
$0.15$0.07
per 1M tokens
Output price
$0.50$0.25
per 1M tokens
Context window
1M
tokens
Hallucination rate
0%
Token Pricing
Token pricing normalized to per-million-token rates.
Input / 1M tokens
$0.15$0.07
Output / 1M tokens
$0.50$0.25
Cache Read / 1M tokens
$0.03$0.01
Token Pricing Details
Rates are shown per 1M tokens for easier comparison.
| Input / 1M tokens | $0.07 |
| Input unit | 1M tokens |
| Output / 1M tokens | $0.25 |
| Output unit | 1M tokens |
| Cache Read / 1M tokens | $0.03$0.01 |
| Cache Read unit | 1M tokens |
Feature Availability
Capabilities explicitly listed in the current payload.
LLM
Available
Vision
Not listed
Function calling
Available
Reasoning
Available
Service tiers
Not listed
Supported Parameters
Code Samples
Quick start with the Routeway API
import OpenAI from 'openai';
const openai = new OpenAI({
baseURL: "https://api.routeway.ai/v1",
apiKey: "<YOUR_API_KEY>",
});
async function main() {
const completion = await openai.chat.completions.create({
model: "glm-5.3-flash",
messages: [
{
role: "user",
content: "Explain quantum computing in simple terms"
}
]
});
console.log(completion.choices[0].message);
}
main();