Llama 3.1 Nemotron 70B Instruct is NVIDIA's language model with a 131K context window and up to 16K output tokens, available from 4 providers, starting at $0.12 / 1M input and $0.3 / 1M output. A 70B instruction-tuned LLM fine-tuned by NVIDIA on Llama 3.1 to significantly improve helpfulness and response quality on user queries.
Specifications
Canonical IDnvidia-llama-3-1-nemotron-70b-instruct
TypeLanguage
StatusDeprecated
CreatorNVIDIANVIDIA
Providers
Context Window131K tokens
Max Output16K tokens
Input ModalitiesText
Output ModalitiesText
Parameters70B
HuggingFace Likes2,064
HuggingFace Downloads (30d)9,963
HuggingFace Downloads (all-time)1,815,075
Release Date · 2 years ago
Knowledge Cutoff · 3 years ago
Deprecation Date
Benchmarks
Intelligence Index
#423
Math Index
#237
MMLU-Pro
#233
GPQA
#418
HLE
#451
LiveCodeBench
#276
AIME
#96
IFBench
#362
Time to First Token
#537
SciCode
#395
MATH-500
#131
AIME 2025
#237
LCR
#388
TerminalBench Hard
#286
TAU2
#307
Output TPS
#56

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning·
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
DeepInfra logo
DeepInfra
deepinfra/nvidia/Llama-3.1-Nemotron-70B-Instruct
$0.6$0.6
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/llama-v3p1-nemotron-70b-instruct
$0.9$0.9
Lambda logo
Lambda
lambda_ai/llama3.1-nemotron-70b-instruct-fp8
$0.12$0.3
Together AI logo
Together AI
together_ai/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
$0.88$0.88

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 3.1 Nemotron 70B Instruct, ranked by cheapest on-demand price. The model needs about 168 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_NP40sAzure4× AMD Alveo U250 FPGA (64GB)256 GB$6.60/hr
g7e.12xlargeAWS2× RTX PRO Server 6000192 GB$8.29/hr
g6e.12xlargeAWS4× L40S179 GB$10.49/hr
7 more instances can run Llama 3.1 Nemotron 70B Instruct
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Llama 3.3 70B Instruct131K$0.100$0.200Deprecated
Llama 3.2 3B Instruct131K$0.015$0.020Deprecated
Llama 3.2 1B Instruct131K$0.020$0.020Deprecated
Llama 3.1 405B Instruct131K$0.120$0.300Deprecated
Llama 3.1 8B Instruct200K$0.020$0.030Deprecated
Llama 3.1 70B Instruct131K$0.100$0.100Deprecated
Llama 3.1 70B128K$0.360$0.360Available
Llama 3.1 8B131K$0.030$0.050Available
Llama 3 70B Instruct131K$0.120$0.300Deprecated
Llama 3 8B Instruct32K$0.030$0.040Deprecated
Llama 3.1 Nemotron 70B Instruct131K$0.120$0.300Current

Model IDs

accounts/fireworks/models/llama-v3p1-nemotron-70b-instruct
deepinfra/nvidia/Llama-3.1-Nemotron-70B-Instruct
fireworks_ai/accounts/fireworks/models/llama-v3p1-nemotron-70b-instruct
lambda_ai/llama3.1-nemotron-70b-instruct-fp8
llama-3-1-nemotron-instruct-70b
nvidia-llama-3-1-nemotron-70b-instruct
nvidia/llama-3.1-nemotron-70b-instruct
nvidia/Llama-3.1-Nemotron-70B-Instruct
together_ai/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF