GLM-4.5 Air is Zhipu AI's language model with a 131K context window and up to 98K output tokens, available from 7 providers, starting at $0.125 / 1M input and $0.45 / 1M output. A compact MoE variant of GLM-4.5 from Z AI, offering a lighter architecture while retaining strong agentic reasoning and tool-use performance.
Specifications
Canonical IDzhipu-glm-4-5-air
TypeLanguage
StatusActive
CreatorZhipu AIZhipu AI
Providers
Context Window131K tokens
Max Output98K tokens
Input ModalitiesText
Output ModalitiesText
Reasoning Effortsdefault
Parameters110B
HuggingFace Likes599
HuggingFace Downloads (30d)389,697
HuggingFace Downloads (all-time)3,025,118
Release Date · 1 year ago
Knowledge Cutoff · 2 years ago
Benchmarks
Intelligence Index
16.7
#223
Math Index
80.7
#60
MMLU-Pro
0.8
#79
GPQA
0.7
#190
HLE
0.1
#227
LiveCodeBench
0.7
#74
AIME
0.7
#38
IFBench
0.4
#266
Time to First Token
0.00s
#377
SciCode
0.3
#248
MATH-500
1.0
#25
AIME 2025
0.8
#60
LCR
0.5
#188
TerminalBench Hard
0.2
#143
TAU2
0.5
#186
Output TPS
0.0
#499

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities7/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/glm-4p5-air
$0.22$0.88N/A
Hugging Face logo
Hugging Face
novita:zai-org/glm-4.5-air
$0.13$0.85N/A
Novita logo
Novita
novita/zai-org/glm-4.5-air
$0.13$0.85N/A
OpenRouter logo
OpenRouter
z-ai/glm-4.5-air
$0.13$0.85$0.025
Pinstripes
pinstripes/ps/glm-4.5-air
$0.125$0.45N/A
Vercel AI Gateway logo
Vercel AI Gateway
zai/glm-4.5-air
$0.2$1.10$0.03
Z AI (Zhipu) logo
Z AI (Zhipu)
zai/glm-4.5-air
$0.2$1.10N/A

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host GLM-4.5 Air, ranked by cheapest on-demand price. The model needs about 265 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g7e.24xlargeAWS4× RTX PRO Server 6000384 GB$16.57/hrus-east-1
g4-standard-192GCP4× nvidia-rtx-pro-6000384 GB$18.00/hrus-central1
p4d.24xlargeAWS8× A100320 GB$21.96/hrus-west-2
7 more instances can run GLM-4.5 Air
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
GLM-4.7205K$0.400$1.75
GLM-4.6205K$0.431$2.00
GLM-4.5131K$0.400$1.60

Model IDs

accounts/fireworks/models/glm-4p5-air
fireworks_ai/accounts/fireworks/models/glm-4p5-air
glm-4-5-air
novita/zai-org/glm-4.5-air
pinstripes/ps/glm-4.5-air
vercel_ai_gateway/zai/glm-4.5-air
z-ai/glm-4.5-air
z-ai/glm-4.5-air:free
zai-org/glm-4.5-air
zai/glm-4.5-air
zhipu-glm-4-5-air