FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling. Inference platform · OpenAI-compatible API · High Throughput · Low Latency · Open Source

Intelligence vs Price

Best value among FriendliAI models on this chart: Llama 3.1 8B Instruct. Prices use each model's lowest available FriendliAI price across regions. Hover any dot for full pricing, or click a creator in the legend to isolate.

Language Models
Intelligence
Blended Price, $
Log X

FriendliAI models

2 models in Global, 2 with pricing
All Model Types
All Creators
US Dollar ($)
Per 1M tokens
Input/1M
to
Output/1M
to
Model
Creator
Input Price, $
Output Price, $
Context
Max Output
Inference Providers
Intelligence
Coding
Llama 3.1 8B InstructMeta logoMeta0.10.1200K128Kcompare (20)7.4#15.4#1
Llama 3.1 70B InstructMeta logoMeta0.60.6131K16Kcompare (12)6.5#2N/A