Llama 4 Scout is Meta's language model with a 1.3M context window and up to 16K output tokens, available from 4 providers, starting at $0.1 / 1M input and $0.3 / 1M output. Meta's Llama 4 Scout MoE LLM with 17B active parameters and 16 experts, offering efficient multimodal inference with native image and text support.
Specifications
Canonical IDmeta-llama-4-scout
TypeLanguage
StatusActive
CreatorMetaMeta
Providers
Context Window1.3M tokens
Max Output16K tokens
Input ModalitiesImageText
Output ModalitiesText
Parameters109B
HuggingFace Likes1,274
HuggingFace Downloads (30d)399,353
HuggingFace Downloads (all-time)5,433,902
Release Date · 1 year ago
Knowledge Cutoff · 2 years ago
Benchmarks
Intelligence Index
10.3
#304
Coding Index
8.2
#153
Math Index
14.0
#214
MMLU-Pro
0.8
#162
GPQA
0.6
#298
HLE
0.0
#423
LiveCodeBench
0.3
#206
AIME
0.3
#79
IFBench
0.4
#241
Time to First Token
0.61s
#420
SciCode
0.2
#393
MATH-500
0.8
#87
AIME 2025
0.1
#214
LCR
0.3
#248
TerminalBench Hard
0.0
#325
TAU2
0.2
#338
Output TPS

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning·
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Input
$ / 1M
Output
$ / 1M
Google Gemini logo
Google Gemini
llama-4-scout
$0.25$0.7$0.125$0.35
Google Vertex AI logo
Google Vertex AI
llama-4-scout
$0.25$0.7$0.125$0.35
OpenRouter logo
OpenRouter
meta-llama/llama-4-scout
$0.1$0.3
Vercel AI Gateway logo
Vercel AI Gateway
meta/llama-4-scout
$0.17$0.66

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 4 Scout, ranked by cheapest on-demand price. The model needs about 261 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g7e.24xlargeAWS4× RTX PRO Server 6000384 GB$16.57/hrus-east-1
g4-standard-192GCP4× nvidia-rtx-pro-6000384 GB$18.00/hrus-central1
p4d.24xlargeAWS8× A100320 GB$21.96/hrus-west-2
7 more instances can run Llama 4 Scout
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
LlamaGuard 4 12B1.0M$0.180$0.180Available
Llama 4 Scout1.3M$0.100$0.300Current
Llama 4 Maverick1.0M$0.120$0.485Available

Model IDs

llama-4-scout
meta-llama-4-scout
meta-llama/llama-4-scout
meta/llama-4-scout
vercel_ai_gateway/meta/llama-4-scout