FriendliAI
FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling. Inference platform · OpenAI-compatible API · High Throughput · Low Latency · Open Source
Intelligence vs Price
Best value among FriendliAI models on this chart: Llama 3.1 8B Instruct. Prices use each model's lowest available FriendliAI price across regions. Hover any dot for full pricing, or click a creator in the legend to isolate.
FriendliAI models
2 models in Global, 2 with pricingModel | Creator | Input Price, $ | Output Price, $ | Context | Max Output | Inference Providers | Intelligence | Coding | |
|---|---|---|---|---|---|---|---|---|---|
| Llama 3.1 8B Instruct | 0.1 | 0.1 | 200K | 128K | compare (20) | 7.4#1 | 5.4#1 | ||
| Llama 3.1 70B Instruct | 0.6 | 0.6 | 131K | 16K | compare (12) | 6.5#2 | N/A |