GLM-5.3 Flash is Zhipu AI's language model with a 1.3M context window and up to 944K output tokens, available from 15 providers, starting at $0.11 / 1M input and $0.35 / 1M output. A native multimodal LLM optimized for efficient coding and long-horizon agent tasks, using a hybrid sparse and linear attention architecture to maintain accuracy over long contexts.
Specifications
Canonical IDzhipu-glm-5-3-flash
TypeLanguage
StatusActive
CreatorZhipu AIZhipu AI
Providers
Context Window1.3M tokens
Max Output944K tokens
Input ModalitiesImageTextVideo
Output ModalitiesText
Reasoning Effortsdefault
Parameters321B
HuggingFace Likes1,107
HuggingFace Downloads (30d)0
HuggingFace Downloads (all-time)0
Release Date · 1 month ago
Deprecation Date
Benchmarks
Intelligence Index
#24
Coding Index
#22
GPQA
#30
HLE
#42
Time to First Token
#529
SciCode
#48
LCR
#48
Output TPS
#204

Capabilities

Input3/5
Text✓
Image✓
Audio·
Video✓
PDF·
Output1/5
Text✓
Image·
Audio·
Video·
Embedding·
Capabilities7/13
Reasoning✓
Adaptive Reasoning·
Function Calling✓
Parallel Function Calling✓
Structured Outputs✓
Native JSON Schema✓
Web Search✓
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching✓
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardFlexFastBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Cache Write 5m
$ / 1M
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Aihubmix
aihubmix/glm-5.3-flash
$0.113$0.394$0.0282N/A—————————
Azure AI Foundry logo
Azure AI Foundry
azure_ai/FW-GLM-5.3-Flash
$0.188$0.625$0.038N/A—————————
Baseten
baseten/zai-org/GLM-5.3-Flash
$0.15$0.5$0.03N/A—————————
Databricks logo
Databricks
databricks/databricks-glm-5-3-flash
$0.15$0.5$0.03$0.15—————————
Fireworks AI logo
Fireworks AI
fireworks_ai/glm-5p3-flash
$0.15$0.5$0.03N/A———$0.188$0.625$0.0375———
FriendliAI logo
FriendliAI
friendliai/zai-org/GLM-5.3-Flash
$0.15$0.5$0.03N/A—————————
Hugging Face logo
Hugging Face
novita:zai-org/glm-5.3-flash
$0.15$0.5N/AN/A—————————
Nebius logo
Nebius
zai-org/GLM-5.3-Flash
$0.15$0.5N/AN/A—————————
OpenRouter logo
OpenRouter
z-ai/glm-5.3-flash
$0.15$0.5$0.03N/A——————N/AN/AN/A
OpenRouter logo
OpenRouter
z-ai/glm-5.3-flash:batch
——————————$0.06$0.2$0.012
Perplexity logo
Perplexity
perplexity/perplexity/glm-5.3-flash
$0.15$0.5$0.03N/A—————————
Sail
sail/zai-org/GLM-5.3-Flash
$0.11$0.35$0.02N/A$0.05$0.18$0.01——————
Together AI logo
Together AI
together_ai/zai-org/GLM-5.3-Flash
$0.15$0.5$0.03N/A—————————
Vercel AI Gateway logo
Vercel AI Gateway
zai/glm-5.3-flash
$0.15$0.5$0.03N/A—————————
Weights & Biases logo
Weights & Biases
wandb/zai-org/GLM-5.3-Flash
$0.15$0.5$0.05N/A—————————
Z AI (Zhipu) logo
Z AI (Zhipu)
zai/glm-5.3-flash
$0.15$0.5$0.03N/A—————————

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host GLM-5.3 Flash, ranked by cheapest on-demand price. The model needs about 771 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_ND96is_MI300X_v5Azure8× AMD Instinct MI300X GPU (192GB)1536 GB$48.00/hr
Standard_ND96isr_MI300X_v5Azure8× AMD Instinct MI300X1536 GB$48.00/hr
p5en.48xlargeAWS8× H2001128 GB$63.30/hr
5 more instances can run GLM-5.3 Flash
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
GLM-5.3 FlashX1.0M$0.370$1.25Available
GLM-5.3 Flash1.3M$0.110$0.350Current
GLM-5.3 Flash US—1.0M$0.225$0.750Available
GLM-4.6V Flash128K——Available
GLM-4.7 Flash Non-Reasoning————Available
GLM-4.5 Flash—128K——Available

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
GLM-5.3 Prime—1.0M$2.80$8.80
GLM-5.3—1.3M$0.070$3.08
GLM-5.3 (50% off)—1.0M——
GLM-5.2 Fast—1.0M$2.10$6.60
GLM-5.2—1.0M$0.171$1.98
GLM-5V Turbo—205K$0.704$3.10
GLM-5 Turbo—203K$1.20$4.00
GLM-5.1 Non-Reasoning—————
GLM-5 Non-Reasoning—————
GLM-5 Code——200K$1.20$5.00

Model IDs

accounts/fireworks/models/glm-5p3-flash
aihubmix/glm-5.3-flash
azure_ai/FW-GLM-5.3-Flash
baseten/zai-org/GLM-5.3-Flash
databricks/databricks-glm-5-3-flash
fireworks_ai/accounts/fireworks/models/glm-5p3-flash
fireworks_ai/glm-5p3-flash
friendliai/zai-org/GLM-5.3-Flash
glm-5-3-flash
nebius/zai-org/GLM-5.3-Flash
openrouter/z-ai/glm-5.3-flash
openrouter/z-ai/glm-5.3-flash:batch
perplexity/perplexity/glm-5.3-flash
sail/zai-org/GLM-5.3-Flash
together_ai/zai-org/GLM-5.3-Flash
wandb/zai-org/GLM-5.3-Flash
z-ai/glm-5.3-flash
zai-org/glm-5.3-flash
zai-org/GLM-5.3-Flash
zai/glm-5.3-flash
zhipu-glm-5-3-flash