FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling. Inference platform · OpenAI-compatible API · High Throughput · Low Latency · Open Source

Intelligence vs Price

Best value among FriendliAI models on this chart: GLM-5.3 · GLM-5.3 Flash. Prices use each model's lowest available FriendliAI price across regions. Hover any dot for full pricing, or click a creator in the legend to isolate.

Language Models
Intelligence
Blended Price, $
Log X

FriendliAI models

8 models in Global, 8 with pricing
All Model Types
All Creators
US Dollar ($)
Per 1M tokens
Input/1M
to
Output/1M
to
Select
Model
Creator
Input Price, $
Output Price, $
Context
Max Output
Inference Providers
Intelligence
Coding
GLM-5.3Zhipu AI logoZhipu AI1.263.961.3M944K#1#1
GLM-5.3 FlashZhipu AI logoZhipu AI0.150.51.3M944K#2#2
GLM-5.2Zhipu AI logoZhipu AI1.404.401.0M944K#3#3
GLM-5.1Zhipu AI logoZhipu AI1.404.40205K203K#4#4
MiniMax M2.5MiniMax logoMiniMax0.31.201.0M131K#5N/A
DeepSeek V3.2DeepSeek logoDeepSeek0.51.50164K147K#6#5
EXAONE 2 750B A37BLG Research logoLG Research0.62.40262KN/AN/AN/A
Gemma 4 31B ITGoogle logoGoogle0.140.4262K131KN/AN/A