GLM-4.6 is Zhipu AI's language model with a 205K context window and up to 131K output tokens, available from 11 providers, starting at $0.43 / 1M input and $1.75 / 1M output. A MoE LLM from Z AI that extends GLM-4.5 with a 200K context window and improved agentic reasoning and tool-use capabilities.
Specifications
Canonical IDzhipu-glm-4-6
TypeLanguage
StatusDeprecated
CreatorZhipu AIZhipu AI
Providers
Context Window205K tokens
Max Output131K tokens
Input ModalitiesText
Output ModalitiesText
Reasoning Effortsdefault
Parameters357B
HuggingFace Likes1,212
HuggingFace Downloads (30d)36,176
HuggingFace Downloads (all-time)681,060
Release Date · 11 months ago
Knowledge Cutoff · 1 year ago
Deprecation Date
Benchmarks
Intelligence Index
#145
Coding Index
#72
Math Index
#57
MMLU-Pro
#67
GPQA
#172
HLE
#150
LiveCodeBench
#88
IFBench
#230
Time to First Token
#491
SciCode
#178
AIME 2025
#57
LCR
#197
TerminalBench Hard
#114
TAU2
#131
Output TPS
#183

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities6/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
$0.431$2.01N/A$0.215$1.00
Baseten
baseten/zai-org/GLM-4.6
$0.6$2.20N/A
Cerebras logo
Cerebras
cerebras/zai-glm-4.6
$2.25$2.75N/A
DeepInfra logo
DeepInfra
deepinfra/zai-org/GLM-4.6
$0.5$2.00$0.1
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/glm-4p6
$0.55$2.19N/A
Hugging Face logo
Hugging Face
novita:zai-org/glm-4.6
$0.55$2.20N/A
Novita logo
Novita
novita/zai-org/glm-4.6
$0.55$2.20$0.11
OpenRouter logo
OpenRouter
z-ai/glm-4.6
$0.43$1.75$0.08
Together AI logo
Together AI
together_ai/zai-org/GLM-4.6
$0.6$2.20N/A
$0.6$2.20$0.11
Z AI (Zhipu) logo
Z AI (Zhipu)
zai/glm-4.6
$0.6$2.20$0.11

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host GLM-4.6, ranked by cheapest on-demand price. The model needs about 856 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_ND96is_MI300X_v5Azure8× AMD Instinct MI300X GPU (192GB)1536 GB$48.00/hr
Standard_ND96isr_MI300X_v5Azure8× AMD Instinct MI300X1536 GB$48.00/hr
p5en.48xlargeAWS8× H2001128 GB$63.30/hr
5 more instances can run GLM-4.6
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
GLM-4.7205K$0.400$1.75Deprecated
GLM-4.6205K$0.430$1.75Current
GLM-4.5131K$0.400$1.60Deprecated
GLM-4.5 Air131K$0.125$0.450Available

Model IDs

accounts/fireworks/models/glm-4p6
baseten/zai-org/GLM-4.6
cerebras/zai-glm-4.6
deepinfra/zai-org/GLM-4.6
dev/glm46
fireworks_ai/accounts/fireworks/models/glm-4p6
glm-4-6
glm-4-6-reasoning
glm-4.6
novita/zai-org/glm-4.6
openrouter/z-ai/glm-4.6
openrouter/z-ai/glm-4.6:exacto
together_ai/zai-org/GLM-4.6
vercel_ai_gateway/zai/glm-4.6
z-ai/glm-4.6
zai-org/glm-4.6
zai-org/GLM-4.6
zai/glm-4.6
zhipu-glm-4-6