Nemotron 3 Ultra 550B A55B is NVIDIA's language model with a 1.0M context window and up to 461K output tokens, available from 6 providers, starting at $0.5 / 1M input and $2.20 / 1M output. A 550B-parameter mixture-of-experts Nemotron model with 55B active parameters, built for frontier-scale reasoning, tool use, and agentic tasks.
Specifications
Canonical IDnvidia-nemotron-3-ultra-550b-a55b
TypeLanguage
StatusDeprecated
CreatorNVIDIANVIDIA
Providers
Context Window1.0M tokens
Max Output461K tokens
Input ModalitiesImageText
Output ModalitiesText
Reasoning Effortsdefault
Parameters550B
HuggingFace Likes225
HuggingFace Downloads (30d)111,067
HuggingFace Downloads (all-time)111,067
Release Date · 3 months ago
Deprecation Date
Benchmarks
Intelligence Index
#72
Coding Index
#65
GPQA
#73
HLE
#78
IFBench
#4
Time to First Token
#452
SciCode
#147
LCR
#76
TerminalBench Hard
#58
TAU2
#107
Output TPS
#53

Capabilities

Input2/5
Text
Image
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities6/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardFreeBatch
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
DeepInfra logo
DeepInfra
deepinfra/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
$0.5$2.20$0.1
Nebius logo
Nebius
nvidia/Nemotron-3-Ultra-550b-a55b
$1.00$3.00N/A
OpenRouter logo
OpenRouter
nvidia/nemotron-3-ultra-550b-a55b
$0.5$2.20$0.1N/AN/AN/AN/AN/A
OpenRouter logo
OpenRouter
nvidia/nemotron-3-ultra-550b-a55b:batch
$0.6$3.60$0.2
OpenRouter logo
OpenRouter
nvidia/nemotron-3-ultra-550b-a55b:free
$N/A$N/AN/AN/AN/A
Together AI logo
Together AI
together_ai/nvidia/nemotron-3-ultra-550b-a55b
$0.6$3.60$0.2
Vercel AI Gateway logo
Vercel AI Gateway
nvidia/nemotron-3-ultra-550b-a55b
$0.6$2.40$0.12
Weights & Biases logo
Weights & Biases
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
$0.75$2.75$0.15

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Nemotron 3 Ultra 550B A55B, ranked by cheapest on-demand price. The model needs about 1320 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_ND96is_MI300X_v5Azure8× AMD Instinct MI300X GPU (192GB)1536 GB$48.00/hr
Standard_ND96isr_MI300X_v5Azure8× AMD Instinct MI300X1536 GB$48.00/hr
p6-b200.48xlargeAWS8× B2001432 GB$113.93/hr
2 more instances can run Nemotron 3 Ultra 550B A55B
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Nemotron 3 Ultra 550B A55B1.0M$0.500$2.20Current
Nemotron 3 Ultra NVfp4262K$0.600$2.40Available

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
Nemotron 4 15B
Nemotron Lightning 3.51.0M$0.050$0.200
Nemotron 3.5 Lightning 30B1.0M
Nemotron 3.5 Content Safety131K$0.200$0.200
Nemotron Nano 3 30B A3B Omni Reasoning256K
Nemotron Super 3 120B256K$0.150$0.650
Nemotron Nano 3 30B262K$0.060$0.240
Nemotron Nano 3 30B A3B Omni
Nemotron Nano 3 30B A3B Reasoning
Nemotron Nano 3 4B

Model IDs

deepinfra/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
huggingface-reasoning-nvidia-nemotron-3-ultra-550b-a55b-nvfp4
nvidia-nemotron-3-ultra-550b-a55b
nvidia/nemotron-3-ultra-550b-a55b
nvidia/nemotron-3-ultra-550b-a55b:batch
nvidia/nemotron-3-ultra-550b-a55b:free
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
together_ai/nvidia/nemotron-3-ultra-550b-a55b
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B