Llama 4 Maverick is Meta's language model with a 1.0M context window and up to 115K output tokens, available from 5 providers, starting at $0.12 / 1M input and $0.485 / 1M output. Meta's Llama 4 Maverick MoE LLM with 128 experts and 17B active parameters, delivering high-capacity multimodal language and vision understanding.
Specifications
Canonical IDmeta-llama-4-maverick
TypeLanguage
StatusActive
CreatorMetaMeta
Providers
Context Window1.0M tokens
Max Output115K tokens
Input ModalitiesImageText
Output ModalitiesText
Parameters402B
HuggingFace Likes478
HuggingFace Downloads (30d)30,421
HuggingFace Downloads (all-time)554,732
Release Date · 1 year ago
Knowledge Cutoff · 2 years ago
Benchmarks
Intelligence Index
#299
Coding Index
#163
Math Index
#220
MMLU-Pro
#109
GPQA
#296
HLE
#387
LiveCodeBench
#192
AIME
#75
IFBench
#236
Time to First Token
#393
SciCode
#293
MATH-500
#81
AIME 2025
#220
LCR
#235
TerminalBench Hard
#253
TAU2
#349
Output TPS
#107

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning·
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Cache Write 5m
$ / 1M
Input
$ / 1M
Output
$ / 1M
Databricks logo
Databricks
databricks/databricks-llama-4-maverick
$0.5$1.50$0.5$0.5
Google Vertex AI logo
Google Vertex AI
llama-4-maverick
$0.35$1.15N/AN/A$0.175$0.575
OpenRouter logo
OpenRouter
meta-llama/llama-4-maverick
$0.188$0.652N/AN/A
Snowflake logo
Snowflake
llama4-maverick
$0.12$0.485N/AN/A$0.06$0.242
Vercel AI Gateway logo
Vercel AI Gateway
meta/llama-4-maverick
$0.24$0.97N/AN/A

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 4 Maverick, ranked by cheapest on-demand price. The model needs about 964 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_ND96is_MI300X_v5Azure8× AMD Instinct MI300X GPU (192GB)1536 GB$48.00/hr
Standard_ND96isr_MI300X_v5Azure8× AMD Instinct MI300X1536 GB$48.00/hr
p5en.48xlargeAWS8× H2001128 GB$63.30/hr
3 more instances can run Llama 4 Maverick
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
LlamaGuard 4 12B1.0M$0.180$0.180Deprecated
Llama 4 Maverick1.0M$0.120$0.485Current
Llama 4 Scout1.3M$0.100$0.300Available

Model IDs

databricks/databricks-llama-4-maverick
llama-4-maverick
meta-llama-4-maverick
meta-llama/llama-4-maverick
meta/llama-4-maverick
openrouter/meta-llama/llama-4-maverick
snowflake/llama4-maverick
vercel_ai_gateway/meta/llama-4-maverick