FriendliAI
FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling. Inference platform · OpenAI-compatible API · High Throughput · Low Latency · Open Source
Intelligence vs Price
Best value among FriendliAI models on this chart: GLM-5.3 · GLM-5.3 Flash · Llama 3.1 8B Instruct. Prices use each model's lowest available FriendliAI price across regions. Hover any dot for full pricing, or click a creator in the legend to isolate.
FriendliAI models
4 models in Global, 4 with pricing| Select | Model | Creator | Input Price, $ | Output Price, $ | Context | Max Output | Inference Providers | Intelligence | Coding |
|---|---|---|---|---|---|---|---|---|---|
| GLM-5.3 | 1.26 | 3.96 | 1.3M | 131K | #1 | #1 | |||
| GLM-5.3 Flash | 0.15 | 0.5 | 1.3M | 131K | #2 | #2 | |||
| Llama 3.1 8B Instruct | 0.1 | 0.1 | 200K | 128K | #3 | #3 | |||
| Llama 3.1 70B Instruct | 0.6 | 0.6 | 131K | 16K | #4 | N/A |