Nemotron 3 Ultra
Provider:NVIDIA
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Key specs
- Context:262.1K
- Max output:16.4K
- Tokenizer:Other
- Released:2026-06-04
Token pricing
- Input price:$0.500 / 1M tokens
- Output price:$2.20 / 1M tokens
- Cache read:$0.100 / 1M tokens
- Blended price:$0.925 / 1M tokens
Modalities
Text
Use Cases
- Deep Reasoning & Analysis:Process long legal docs, academic papers, financial reports; multi-hop reasoning and evidence-chain tracing across huge contexts.(≈ $0.018/call)
FAQ
- What is the token pricing of Nemotron 3 Ultra? Input price $0.500 / 1M tokens, output price $2.20 / 1M tokens, blended about $0.925 / 1M tokens. Refer to NVIDIA’s official page for the exact rate.
- What context window does Nemotron 3 Ultra support? Context window is 262.1K, max output about 16.4K, suitable for long documents, multi-turn chat and complex reasoning.
- What scenarios is Nemotron 3 Ultra best for? Based on its capability and pricing, Nemotron 3 Ultra fits best: Deep Reasoning & Analysis. See the Use Cases section for details and cost estimates.
- What input/output modalities does Nemotron 3 Ultra support? Supports Text modality.
- Is Nemotron 3 Ultra a free model? Nemotron 3 Ultra is billed per token, not a free model.
- Which provider offers Nemotron 3 Ultra? Nemotron 3 Ultra is offered by NVIDIA.
More models from NVIDIA
Full info:Nemotron 3 Ultra