Back to Models
Zhipu

GLM 5.3 Flash

Available
50% offA 50% discount applies to input, output, and cached token pricing.

Description

GLM-5.3-Flash is a native multimodal model from Z.ai, optimized for efficient coding and long-horizon agent tasks. Built on a 320B-parameter MoE architecture with 18B active parameters, it supports text, image, and video input with a hybrid sparse and linear attention design that maintains accurate long-context behavior while reducing compute overhead. Despite its agent-first positioning, it is surprisingly good at roleplay.

At a Glance

Key pricing and model details available for this model.

Input price

$0.15$0.07

per 1M tokens

Output price

$0.50$0.25

per 1M tokens

Context window

1M

tokens

Hallucination rate

0%

Token Pricing

Token pricing normalized to per-million-token rates.

Input / 1M tokens

$0.15$0.07

Output / 1M tokens

$0.50$0.25

Cache Read / 1M tokens

$0.03$0.01

Token Pricing Details

Rates are shown per 1M tokens for easier comparison.

Input / 1M tokens$0.07
Input unit1M tokens
Output / 1M tokens$0.25
Output unit1M tokens
Cache Read / 1M tokens$0.03$0.01
Cache Read unit1M tokens

Feature Availability

Capabilities explicitly listed in the current payload.

LLM

Available

Yes

Vision

Not listed

No

Function calling

Available

Yes

Reasoning

Available

Yes

Service tiers

Not listed

No

Supported Parameters

frequency_penalty
logit_bias
max_completion_tokens
presence_penalty
response_format
stop
temperature
tool_choice
tools
top_p

Code Samples

Quick start with the Routeway API

import OpenAI from 'openai';

const openai = new OpenAI({
  baseURL: "https://api.routeway.ai/v1",
  apiKey: "<YOUR_API_KEY>",
});

async function main() {
  const completion = await openai.chat.completions.create({
    model: "glm-5.3-flash",
    messages: [
      {
        role: "user",
        content: "Explain quantum computing in simple terms"
      }
    ]
  });

  console.log(completion.choices[0].message);
}

main();