GLM-4.5 is Zhipu AI's language model with a 131K context window and up to 98K output tokens, available from 7 providers, starting at $0.4 / 1M input and $1.60 / 1M output. A 355B MoE foundation LLM from Z AI with 32B active parameters, designed for intelligent agents with strong reasoning and tool-use capabilities.
Specifications
Canonical IDzhipu-glm-4-5
TypeLanguage
StatusDeprecated
CreatorZhipu AIZhipu AI
Providers
Context Window131K tokens
Max Output98K tokens
Input ModalitiesText
Output ModalitiesText
Reasoning Effortsdefault
Parameters358B
HuggingFace Likes1,398
HuggingFace Downloads (30d)70,876
HuggingFace Downloads (all-time)400,488
Release Date · 1 year ago
Knowledge Cutoff · 2 years ago
Deprecation Date
Benchmarks
Intelligence Index
#218
Math Index
#109
MMLU-Pro
#58
GPQA
#170
HLE
#165
LiveCodeBench
#59
AIME
#14
IFBench
#225
Time to First Token
#312
SciCode
#245
MATH-500
#24
AIME 2025
#109
LCR
#211
TerminalBench Hard
#145
TAU2
#209
Output TPS
#542

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities5/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
DeepInfra logo
DeepInfra
deepinfra/zai-org/GLM-4.5
$0.4$1.60N/A
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/glm-4p5
$0.55$2.19N/A
Novita logo
Novita
novita/zai-org/glm-4.5
$0.6$2.20$0.11
OpenRouter logo
OpenRouter
z-ai/glm-4.5
$0.6$2.20$0.11
$0.6$2.20$0.11
Weights & Biases logo
Weights & Biases
wandb/zai-org/GLM-4.5
$55.00$200.00N/A
Z AI (Zhipu) logo
Z AI (Zhipu)
zai/glm-4.5
$0.6$2.20N/A

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host GLM-4.5, ranked by cheapest on-demand price. The model needs about 860 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_ND96is_MI300X_v5Azure8× AMD Instinct MI300X GPU (192GB)1536 GB$48.00/hr
Standard_ND96isr_MI300X_v5Azure8× AMD Instinct MI300X1536 GB$48.00/hr
p5en.48xlargeAWS8× H2001128 GB$63.30/hr
5 more instances can run GLM-4.5
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
GLM-4.7205K$0.400$1.75Deprecated
GLM-4.6205K$0.430$1.75Deprecated
GLM-4.5131K$0.400$1.60Current
GLM-4.5 Air131K$0.125$0.450Available

Model IDs

accounts/fireworks/models/glm-4p5
deepinfra/zai-org/GLM-4.5
fireworks_ai/accounts/fireworks/models/glm-4p5
glm-4.5
novita/zai-org/glm-4.5
vercel_ai_gateway/zai/glm-4.5
wandb/zai-org/GLM-4.5
z-ai/glm-4.5
zai-org/glm-4.5
zai/glm-4.5
zhipu-glm-4-5