Cerebras
Cerebras powers the world's fastest AI inference on the biggest wafer chip. Cerebras CS-4 delivers up to 30x faster inference than GPUs. Inference platform · OpenAI-compatible API · Fast Inference · Open Weight · Ultra Fast · Wafer Scale
Intelligence vs Price
Best value among Cerebras models on this chart: Qwen3.8 27B · GPT OSS 120B. Prices use each model's lowest available Cerebras price across regions. Hover any dot for full pricing, or click a creator in the legend to isolate.
Cerebras models
7 models in Global, 7 with pricing| Select | Model | Creator | Input Price, $ | Output Price, $ | Context | Max Output | Inference Providers | Intelligence | Coding |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.8 27B | 0.99 | 1.49 | 1.0M | 236K | #1 | #1 | |||
| Gemma 4 31B | 0.99 | 1.49 | 256K | 41K | #2 | #2 | |||
| GPT OSS 120B | 0.35 | 0.75 | 131K | 131K | #3 | #3 | |||
| Qwen3 32B | 0.4 | 0.8 | 131K | 41K | #4 | #4 | |||
| Llama 3.1 70B | 0.6 | 0.6 | 128K | 16K | N/A | N/A | |||
| Llama 3.1 8B | 0.1 | 0.1 | 131K | 16K | N/A | N/A | |||
| Llama 3.3 70B | 0.85 | 1.20 | 128K | 16K | N/A | N/A |