Nemotron 3 Ultra NVfp4 is NVIDIA's language model with a 262K context window, starting at $0.6 / 1M input and $2.40 / 1M output. A large-scale Nemotron LLM optimized for NVIDIA's NVfp4 quantization format, targeting high-throughput inference on NVIDIA hardware.
Specifications
Canonical IDnvidia-nemotron-3-ultra-nvfp4
TypeLanguage
StatusActive
CreatorNVIDIANVIDIA
Providers
Context Window262K tokens
Input ModalitiesText
Output ModalitiesText
Reasoning Effortsdefault

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities3/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Azure AI Foundry logo
Azure AI Foundry
azure_ai/FW-Nemotron-3-Ultra-NVFP4
$0.6$2.40$0.119

Cost Calculator

US Dollar ($)
Preset:

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Nemotron 3 Ultra 550B A55B1.0M$0.600$2.40Available
Nemotron 3 Ultra NVfp4262K$0.600$2.40Current

Other Models

ModelTierReleasedContextInput / 1MOutput / 1M
Nemotron 4 15B
Nemotron Lightning 3.51.0M$0.050$0.200
Nemotron 3.5 Lightning 30B1.0M
Nemotron 3.5 Content Safety128K
Nemotron Nano 3 30B A3B Omni Reasoning256K
Nemotron Super 3 120B256K$0.150$0.650
Nemotron Nano 3 30B262K$0.060$0.240
Nemotron Nano 3 30B A3B Omni
Nemotron Nano 3 30B A3B Reasoning
Nemotron Nano 3 4B

Model IDs

azure_ai/FW-Nemotron-3-Ultra-NVFP4
nvidia-nemotron-3-ultra-nvfp4