Llama 4 Maverick is Meta's language model with a 1.0M context window and up to 115K output tokens, available from 5 providers, starting at $0.12 / 1M input and $0.485 / 1M output. Meta's Llama 4 Maverick MoE LLM with 128 experts and 17B active parameters, delivering high-capacity multimodal language and vision understanding.
Specifications
Canonical IDmeta-llama-4-maverick
TypeLanguage
StatusActive
CreatorMetaMeta
Providers
Context Window1.0M tokens
Max Output115K tokens
Input ModalitiesImageText
Output ModalitiesText
Parameters402B
HuggingFace Likes478
HuggingFace Downloads (30d)30,421
HuggingFace Downloads (all-time)554,732
Release Date · 1 year ago
Knowledge Cutoff · 2 years ago
Benchmarks
Intelligence Index
#279
Coding Index
#157
Math Index
#220
MMLU-Pro
#109
GPQA
#287
HLE
#373
LiveCodeBench
#192
AIME
#75
IFBench
#236
Time to First Token
#373
SciCode
#269
MATH-500
#81
AIME 2025
#220
LCR
#217
TerminalBench Hard
#252
TAU2
#348
Output TPS
#150

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning·
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Cache Write 5m
$ / 1M
Input
$ / 1M
Output
$ / 1M
Databricks logo
Databricks
databricks/databricks-llama-4-maverick
$0.5$1.50$0.5$0.5
Google Vertex AI logo
Google Vertex AI
llama-4-maverick
$0.35$1.15N/AN/A$0.175$0.575
OpenRouter logo
OpenRouter
meta-llama/llama-4-maverick
$0.2$0.696N/AN/A
Snowflake logo
Snowflake
llama4-maverick
$0.12$0.485N/AN/A$0.06$0.242
Vercel AI Gateway logo
Vercel AI Gateway
meta/llama-4-maverick
$0.24$0.97N/AN/A

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 4 Maverick, ranked by cheapest on-demand price. The model needs about 964 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_ND96is_MI300X_v5Azure8× AMD Instinct MI300X GPU (192GB)1536 GB$48.00/hr
Standard_ND96isr_MI300X_v5Azure8× AMD Instinct MI300X1536 GB$48.00/hr
p5en.48xlargeAWS8× H2001128 GB$63.30/hr
5 more instances can run Llama 4 Maverick
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
LlamaGuard 4 12B1.0M$0.180$0.180Deprecated
Llama 4 Maverick1.0M$0.120$0.485Current
Llama 4 Scout1.3M$0.100$0.300Available

Model IDs

databricks/databricks-llama-4-maverick
llama-4-maverick
meta-llama-4-maverick
meta-llama/llama-4-maverick
meta/llama-4-maverick
snowflake/llama4-maverick
vercel_ai_gateway/meta/llama-4-maverick