GLM-4.7 Flash is Zhipu AI's language model with a 203K context window and up to 131K output tokens, available from 6 providers, starting at $0.06 / 1M input and $0.4 / 1M output. A lightweight 30B-A3B MoE model from Z AI that balances strong performance with efficiency, optimized for fast inference and agentic tasks.
Specifications
Canonical IDzhipu-glm-4-7-flash
TypeLanguage
StatusActive
CreatorZhipu AIZhipu AI
Providers
Context Window203K tokens
Max Output131K tokens
Input ModalitiesImageText
Output ModalitiesText
Reasoning Effortsdefault
Parameters31.2B
HuggingFace Likes1,708
HuggingFace Downloads (30d)682,370
HuggingFace Downloads (all-time)4,269,790
Release Date · 9 months ago
Deprecation Date
Benchmarks
Intelligence Index
#223
GPQA
#355
HLE
#281
IFBench
#115
Time to First Token
#444
SciCode
#286
LCR
#275
TerminalBench Hard
#147
TAU2
#4
Output TPS
#103

Capabilities

Input2/5
Text✓
Image✓
Audio·
Video·
PDF·
Output1/5
Text✓
Image·
Audio·
Video·
Embedding·
Capabilities5/13
Reasoning✓
Adaptive Reasoning·
Function Calling✓
Parallel Function Calling·
Structured Outputs✓
Native JSON Schema✓
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching✓
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardFlexFastBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Input
$ / 1M
Output
$ / 1M
Input
$ / 1M
Output
$ / 1M
Amazon Bedrock logo
Amazon Bedrock
zai.glm-4.7-flash
$0.07$0.4N/A$0.035$0.2$0.122$0.7$0.035$0.2
DeepInfra logo
DeepInfra
deepinfra/zai-org/GLM-4.7-Flash
$0.06$0.4$0.01——————
Hugging Face logo
Hugging Face
novita:zai-org/glm-4.7-flash
$0.07$0.4N/A——————
Novita logo
Novita
novita/zai-org/glm-4.7-flash
$0.07$0.4$0.01——————
OpenRouter logo
OpenRouter
z-ai/glm-4.7-flash
$0.0605$0.4$0.01——————
Vercel AI Gateway logo
Vercel AI Gateway
zai/glm-4.7-flash
$0.07$0.4N/A——————

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host GLM-4.7 Flash, ranked by cheapest on-demand price. The model needs about 75 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_NP20sAzure2× AMD Alveo U250 FPGA (64GB)128 GB$3.30/hr
g7e.2xlargeAWSRTX PRO Server 600096 GB$3.36/hr
Standard_NC24ads_A100_v4AzureNVIDIA A10080 GB$3.67/hr
7 more instances can run GLM-4.7 Flash
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Model IDs

accounts/fireworks/models/glm-4p7-flash
cloudflare/@cf/zai-org/glm-4.7-flash
deepinfra/zai-org/GLM-4.7-Flash
glm-4-7-flash
glm-4-7-flash-non-reasoning
novita/zai-org/glm-4.7-flash
openrouter/z-ai/glm-4.7-flash
z-ai/glm-4.7-flash
zai-org/glm-4.7-flash
zai-org/GLM-4.7-Flash
zai.glm-4.7-flash
zai/glm-4.7-flash
zhipu-glm-4-7-flash