Llama 3.3 Nemotron Super 49B is NVIDIA's language model. A 49B-parameter compute-efficient LLM fine-tuned by NVIDIA on Llama 3.3, targeting multi-agent and agentic system workloads with the Nemotron Super architecture.
Specifications
Canonical IDnvidia-llama-3-3-nemotron-super-49b
TypeLanguage
StatusActive
CreatorNVIDIANVIDIA
Input ModalitiesText
Output ModalitiesText
Parameters49B
Benchmarks
Intelligence Index
#316
Math Index
#158
MMLU-Pro
#153
GPQA
#305
HLE
#322
LiveCodeBench
#235
AIME
#60
IFBench
#268
Time to First Token
#222
SciCode
#323
MATH-500
#43
AIME 2025
#158
LCR
#338
TerminalBench Hard
#405
TAU2
#279
Output TPS
#456

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 3.3 Nemotron Super 49B, ranked by cheapest on-demand price. The model needs about 118 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_NP20sAzure2× AMD Alveo U250 FPGA (64GB)128 GB$3.30/hr
Standard_NP40sAzure4× AMD Alveo U250 FPGA (64GB)256 GB$6.60/hr
Standard_NC48ads_A100_v4Azure2× NVIDIA A100160 GB$7.35/hr
7 more instances can run Llama 3.3 Nemotron Super 49B
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Llama Nemotron 1.5 Super 49BAvailable
Llama Nemotron 1.5 Super 49B ReasoningAvailable
Llama 3.3 Nemotron Super 49BCurrent
Llama 3.3 Nemotron Super 49B ReasoningAvailable

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
Llama 3.3 70B Instruct131K$0.120$0.200
Llama 3.2 3B Instruct131K$0.015$0.020
Llama 3.2 1B Instruct131K$0.020$0.020
Llama 3.2 11B128K$0.160$0.160
Llama 3.1 405B Instruct131K$0.120$0.300
Llama 3.1 8B Instruct200K$0.020$0.030
Llama 3.1 70B Instruct131K$0.100$0.100
Llama 3.1 70B128K$0.360$0.360
Llama 3.1 8B131K$0.030$0.050
Llama 3 70B Instruct131K$0.120$0.300

Model IDs

llama-3-3-nemotron-super-49b
llama-3-3-nemotron-super-49b-reasoning
nvidia-llama-3-3-nemotron-super-49b