Llama 4 Scout is Meta's language model with a 1.3M context window and up to 16K output tokens, available from 3 providers, starting at $0.11 / 1M input and $0.34 / 1M output. Meta's Llama 4 Scout MoE LLM with 17B active parameters and 16 experts, offering efficient multimodal inference with native image and text support.
Specifications
Canonical IDmeta-llama-4-scout
TypeLanguage
StatusActive
CreatorMetaMeta
Providers
Context Window1.3M tokens
Max Output16K tokens
Input ModalitiesImageText
Output ModalitiesText
Parameters109B
HuggingFace Likes1,274
HuggingFace Downloads (30d)399,353
HuggingFace Downloads (all-time)5,433,902
Release Date · 1 year ago
Knowledge Cutoff · 2 years ago
Benchmarks
Intelligence Index
#347
Coding Index
#189
Math Index
#231
MMLU-Pro
#186
GPQA
#343
HLE
#472
LiveCodeBench
#222
AIME
#90
IFBench
#267
Time to First Token
#350
SciCode
#429
MATH-500
#101
AIME 2025
#231
LCR
#289
TerminalBench Hard
#345
TAU2
#361
Output TPS
#75

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning·
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Google Vertex AI logo
Google Vertex AI
llama-4-scout
$0.25$0.7N/A$0.125$0.35
OpenRouter logo
OpenRouter
meta-llama/llama-4-scout
$0.11$0.34$0.055
Vercel AI Gateway logo
Vercel AI Gateway
meta/llama-4-scout
$0.17$0.66N/A

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 4 Scout, ranked by cheapest on-demand price. The model needs about 261 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_NC96ads_A100_v4Azure4× NVIDIA A100320 GB$14.69/hr
g7e.24xlargeAWS4× RTX PRO Server 6000384 GB$16.57/hr
g4-standard-192GCP4× nvidia-rtx-pro-6000384 GB$18.00/hr
7 more instances can run Llama 4 Scout
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
LlamaGuard 4 12B1.0M$0.180$0.180Deprecated
Llama 4 Scout1.3M$0.110$0.340Current
Llama 4 Maverick1.0M$0.120$0.485Available

Model IDs

llama-4-scout
meta-llama-4-scout
meta-llama/llama-4-scout
meta/llama-4-scout
vercel_ai_gateway/meta/llama-4-scout