GLM-4.5V is Zhipu AI's language model with a 131K context window and up to 32K output tokens, available from 6 providers, starting at $0.6 / 1M input and $1.20 / 1M output. A multimodal MoE vision-language model from Z AI based on GLM-4.5 Air, delivering strong visual reasoning and tool-use performance.
Specifications
Canonical IDzhipu-glm-4-5v
TypeLanguage
StatusDeprecating
CreatorZhipu AIZhipu AI
Providers
Context Window131K tokens
Max Output32K tokens
Input ModalitiesImageText
Output ModalitiesText
Reasoning Effortsdefault
Parameters108B
HuggingFace Likes717
HuggingFace Downloads (30d)44,600
HuggingFace Downloads (all-time)417,587
Release Date · 1 year ago
Knowledge Cutoff · 2 years ago
Deprecation Date
Benchmarks
Intelligence Index
6.8
#377
Math Index
15.3
#210
MMLU-Pro
0.8
#165
GPQA
0.6
#306
HLE
0.0
#455
LiveCodeBench
0.4
#184
IFBench
0.3
#349
Time to First Token
0.00s
#378
SciCode
0.2
#378
AIME 2025
0.2
#210
LCR
0.0
#425
TerminalBench Hard
0.1
#241
TAU2
0.2
#318
Output TPS
0.0
#500

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities5/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/glm-4p5v
$1.20$1.20N/A
Hugging Face logo
Hugging Face
novita:zai-org/glm-4.5v
$0.6$1.80N/A
Novita logo
Novita
novita/zai-org/glm-4.5v
$0.6$1.80$0.11
OpenRouter logo
OpenRouter
z-ai/glm-4.5v
$0.6$1.80$0.11
Vercel AI Gateway logo
Vercel AI Gateway
zai/glm-4.5v
$0.6$1.80$0.11
Z AI (Zhipu) logo
Z AI (Zhipu)
zai/glm-4.5v
$0.6$1.80N/A

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host GLM-4.5V, ranked by cheapest on-demand price. The model needs about 259 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g7e.24xlargeAWS4× RTX PRO Server 6000384 GB$16.57/hrus-east-1
g4-standard-192GCP4× nvidia-rtx-pro-6000384 GB$18.00/hrus-central1
p4d.24xlargeAWS8× A100320 GB$21.96/hrus-west-2
7 more instances can run GLM-4.5V
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
GLM-4.6V131K$0.300$0.900Available
GLM-4.5V131K$0.600$1.20Current

Model IDs

accounts/fireworks/models/glm-4p5v
fireworks_ai/accounts/fireworks/models/glm-4p5v
glm-4-5v
glm-4-5v-reasoning
novita/zai-org/glm-4.5v
z-ai/glm-4.5v
zai-org/glm-4.5v
zai/glm-4.5v
zhipu-glm-4-5v