Cerebras powers the world's fastest AI inference on the biggest wafer chip. Cerebras CS-4 delivers up to 30x faster inference than GPUs. Inference platform · OpenAI-compatible API · Fast Inference · Open Weight · Ultra Fast · Wafer Scale

Intelligence vs Price

Best value among Cerebras models on this chart: GLM-4.7 · Gemma 4 31B · GPT OSS 120B. Prices use each model's lowest available Cerebras price across regions. Hover any dot for full pricing, or click a creator in the legend to isolate.

Language Models
Intelligence
Blended Price, $
Log X

Cerebras models

8 models in Global, 8 with pricing
All Model Types
All Creators
US Dollar ($)
Per 1M tokens
Input/1M
to
Output/1M
to
Select
Model
Creator
Input Price, $
Output Price, $
Context
Max Output
Inference Providers
Intelligence
Coding
GLM-4.7Zhipu AI logoZhipu AI2.252.75205K131K#1#2
Gemma 4 31BGoogle logoGoogle0.991.49256K41K#2#3
GLM-4.6Zhipu AI logoZhipu AI2.252.75205K131K#3#1
GPT OSS 120BOpenAI logoOpenAI0.350.75131K131K#4#4
Qwen3 32BAlibaba logoAlibaba0.40.8131K41K#5#5
Llama 3.1 70BMeta logoMeta0.60.6128K16KN/AN/A
Llama 3.1 8BMeta logoMeta0.10.1131K16KN/AN/A
Llama 3.3 70BMeta logoMeta0.851.20128K16KN/AN/A