Llama 3.2 1B Instruct is Meta's language model with a 131K context window and up to 54K output tokens, available from 7 providers, starting at $0.02 / 1M input and $0.02 / 1M output. Meta's 1B instruction-tuned LLM from Llama 3.2, optimized for lightweight on-device deployment with efficient instruction-following in multiple regions.
Specifications
Canonical IDmeta-llama-3-2-1b-instruct
TypeLanguage
StatusDeprecated
CreatorMetaMeta
Providers
Context Window131K tokens
Max Output54K tokens
Input ModalitiesText
Output ModalitiesText
Parameters1B
HuggingFace Likes1,372
HuggingFace Downloads (30d)4,623,641
HuggingFace Downloads (all-time)59,725,960
Release Date · 2 years ago
Knowledge Cutoff · 3 years ago
Deprecation Date
Benchmarks
Intelligence Index
#543
Math Index
#273
MMLU-Pro
#336
GPQA
#525
HLE
#329
LiveCodeBench
#333
AIME
#179
IFBench
#403
Time to First Token
#192
SciCode
#513
MATH-500
#191
AIME 2025
#273
LCR
#387
TerminalBench Hard
#399
TAU2
#407
Output TPS
#427

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning·
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Amazon Bedrock logo
Amazon Bedrock
meta.llama3-2-1b-instruct-v1:0
$0.1$0.1
Cloudflare Workers AI logo
Cloudflare Workers AI
@cf/meta/llama-3.2-1b-instruct
$0.027$0.201
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/llama-v3p2-1b-instruct
$0.1$0.1
IBM watsonx logo
IBM watsonx
watsonx/meta-llama/llama-3-2-1b-instruct
$0.1$0.1
Novita logo
Novita
novita/meta-llama/llama-3.2-1b-instruct
$0.02$0.02
OpenRouter logo
OpenRouter
meta-llama/llama-3.2-1b-instruct
$0.027$0.201
SambaNova logo
SambaNova
sambanova/Meta-Llama-3.2-1B-Instruct
$0.04$0.08

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 3.2 1B Instruct, ranked by cheapest on-demand price. The model needs about 2 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g6f.largeAWSL43 GB$0.202/hr
Standard_NV4as_v4AzureAMD Radeon Instinct MI2516 GB$0.233/hr
g6f.xlargeAWSL43 GB$0.237/hr
7 more instances can run Llama 3.2 1B Instruct
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Llama 3.3 70B Instruct131K$0.120$0.200Deprecated
Llama 3.2 1B Instruct131K$0.020$0.020Current
Llama 3.2 3B Instruct131K$0.015$0.020Deprecated
Llama 3.2 11B128K$0.160$0.160Available
Llama 3.1 405B Instruct131K$0.120$0.300Deprecated
Llama 3.1 8B Instruct200K$0.020$0.030Deprecated
Llama 3.1 70B Instruct131K$0.100$0.100Deprecated
Llama 3.1 70B128K$0.360$0.360Available
Llama 3.1 8B131K$0.030$0.050Available
Llama 3 70B Instruct131K$0.120$0.300Deprecated
Llama 3 8B Instruct32K$0.030$0.040Available

Model IDs

@cf/meta/llama-3.2-1b-instruct
accounts/fireworks/models/llama-v3p2-1b-instruct
cloudflare/@cf/meta/llama-3.2-1b-instruct
eu.meta.llama3-2-1b-instruct-v1:0
fireworks_ai/accounts/fireworks/models/llama-v3p2-1b-instruct
llama-3-2-instruct-1b
meta-llama-3-2-1b-instruct
meta-llama/llama-3.2-1b-instruct
meta-textgeneration-llama-3-2-1b-instruct
meta-textgenerationneuron-llama-3-2-1b-instruct
meta.llama3-2-1b-instruct-v1:0
meta.llama3-2-1b-instruct-v1:0:128k
novita/meta-llama/llama-3.2-1b-instruct
sambanova/Meta-Llama-3.2-1B-Instruct
us.meta.llama3-2-1b-instruct-v1:0
watsonx/meta-llama/llama-3-2-1b-instruct