GLM-5.3 Flash is Zhipu AI's language model with a 1.3M context window and up to 131K output tokens, available from 3 providers, starting at $0.075 / 1M input and $0.25 / 1M output. A native multimodal LLM optimized for efficient coding and long-horizon agent tasks, using a hybrid sparse and linear attention architecture to maintain accuracy over long contexts.
Specifications
Canonical IDzhipu-glm-5-3-flash
TypeLanguage
StatusActive
CreatorZhipu AIZhipu AI
Providers
Context Window1.3M tokens
Max Output131K tokens
Input ModalitiesImageTextVideo
Output ModalitiesText
Reasoning Effortsdefault
Parameters321B
HuggingFace Likes1,107
HuggingFace Downloads (30d)0
HuggingFace Downloads (all-time)0
Release Date · 2 days ago
Deprecation Date
Benchmarks
Intelligence Index
#9
Coding Index
#17
GPQA
#24
HLE
#30
Time to First Token
#419
SciCode
#63
LCR
#19
Output TPS
#206

Capabilities

Input3/5
Text
Image
Audio·
Video
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities5/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Hugging Face logo
Hugging Face
together_ai:zai-org/GLM-5.3-Flash
$0.15$0.5N/A
OpenRouter logo
OpenRouter
z-ai/glm-5.3-flash
$0.075$0.25$0.015
Vercel AI Gateway logo
Vercel AI Gateway
zai/glm-5.3-flash
$0.15$0.5$0.03

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host GLM-5.3 Flash, ranked by cheapest on-demand price. The model needs about 771 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_ND96is_MI300X_v5Azure8× AMD Instinct MI300X GPU (192GB)1536 GB$48.00/hr
Standard_ND96isr_MI300X_v5Azure8× AMD Instinct MI300X1536 GB$48.00/hr
p5en.48xlargeAWS8× H2001128 GB$63.30/hr
5 more instances can run GLM-5.3 Flash
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
GLM-5.3 Flash1.3M$0.075$0.250Current
GLM-4.6V Flash128KAvailable
GLM-4.7 Flash Non-ReasoningAvailable
GLM-4.5 Flash128KAvailable

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
GLM-5.31.0M$1.40$4.40
GLM-5.2 Fast1.0M$2.10$6.60
GLM-5.21.0M$0.610$1.98
GLM-5V Turbo203K$1.20$4.00
GLM-5 Turbo203K$1.20$4.00
GLM-5.1 Non-Reasoning
GLM-5 Non-Reasoning
GLM-5 Code200K$1.20$5.00
GLM-5.1 Fast203K$2.80$8.80
GLM-5.1 NVFP4 MTP203K$1.40$4.40

Model IDs

glm-5-3-flash
z-ai/glm-5.3-flash
zai-org/glm-5.3-flash
zai-org/GLM-5.3-Flash
zai/glm-5.3-flash
zhipu-glm-5-3-flash