FriendliAI
FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling. Inference platform · OpenAI-compatible API · High Throughput · Low Latency · Open Source
Intelligence vs Price
Best value among FriendliAI models on this chart: GLM-5.3 · GLM-5.3 Flash. Prices use each model's lowest available FriendliAI price across regions. Hover any dot for full pricing, or click a creator in the legend to isolate.
FriendliAI models
8 models in Global, 8 with pricing| Select | Model | Creator | Input Price, $ | Output Price, $ | Context | Max Output | Inference Providers | Intelligence | Coding |
|---|---|---|---|---|---|---|---|---|---|
| GLM-5.3 | 1.26 | 3.96 | 1.3M | 944K | #1 | #1 | |||
| GLM-5.3 Flash | 0.15 | 0.5 | 1.3M | 944K | #2 | #2 | |||
| GLM-5.2 | 1.40 | 4.40 | 1.0M | 944K | #3 | #3 | |||
| GLM-5.1 | 1.40 | 4.40 | 205K | 203K | #4 | #4 | |||
| MiniMax M2.5 | 0.3 | 1.20 | 1.0M | 131K | #5 | N/A | |||
| DeepSeek V3.2 | 0.5 | 1.50 | 164K | 147K | #6 | #5 | |||
| EXAONE 2 750B A37B | 0.6 | 2.40 | 262K | N/A | N/A | N/A | |||
| Gemma 4 31B IT | 0.14 | 0.4 | 262K | 131K | N/A | N/A |