Cerebras
Cerebras powers the world's fastest AI inference on the biggest wafer chip. Cerebras CS-4 delivers up to 30x faster inference than GPUs. Inference platform · OpenAI-compatible API · Fast Inference · Open Weight · Ultra Fast · Wafer Scale
Intelligence vs Price
Best value among Cerebras models on this chart: GLM-4.7 · Gemma 4 31B · GPT OSS 120B. Prices use each model's lowest available Cerebras price across regions. Hover any dot for full pricing, or click a creator in the legend to isolate.
Cerebras models
8 models in Global, 8 with pricing| Select | Model | Creator | Input Price, $ | Output Price, $ | Context | Max Output | Inference Providers | Intelligence | Coding |
|---|---|---|---|---|---|---|---|---|---|
| GLM-4.7 | 2.25 | 2.75 | 205K | 131K | #1 | #2 | |||
| Gemma 4 31B | 0.99 | 1.49 | 256K | 41K | #2 | #3 | |||
| GLM-4.6 | 2.25 | 2.75 | 205K | 131K | #3 | #1 | |||
| GPT OSS 120B | 0.35 | 0.75 | 131K | 131K | #4 | #4 | |||
| Qwen3 32B | 0.4 | 0.8 | 131K | 41K | #5 | #5 | |||
| Llama 3.1 70B | 0.6 | 0.6 | 128K | 16K | N/A | N/A | |||
| Llama 3.1 8B | 0.1 | 0.1 | 131K | 16K | N/A | N/A | |||
| Llama 3.3 70B | 0.85 | 1.20 | 128K | 16K | N/A | N/A |