Nemotron 3 Ultra (batch)
Nemotron 3 Ultra · Batch tier (Batch)
Asynchronous batch tier: discounted offline batch requests with high throughput and relaxed latency — ideal for large-scale non-real-time workloads.
Best for: Large-scale offline inference, cost-sensitive batch jobs, backfills
This page covers the Nemotron 3 Ultra Batch tier, sharing the underlying model capability with its real-time standard tier; the main difference is in billing mode and invocation method. See the Nemotron 3 Ultra standard real-time tier.
Provider:NVIDIA
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Key specs
- Context:512.3K
- Max output:461.1K
- Tokenizer:Other
- Released:2026-06-04
Token pricing
- Input price:$0.600 / 1M tokens
- Output price:$3.60 / 1M tokens
- Cache read:$0.200 / 1M tokens
- Blended price:$1.35 / 1M tokens
Modalities
Text
Use Cases
- Deep Reasoning & Analysis:Process long legal docs, academic papers, financial reports; multi-hop reasoning and evidence-chain tracing across huge contexts.(≈ $0.029/call)
When to pick this model
Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.
- Long max output (>=30K tokens): strong fit for deep report / essay / document-drafting workflows.
FAQ
- What is the token pricing of Nemotron 3 Ultra (batch)? Input price $0.600 / 1M tokens, output price $3.60 / 1M tokens, blended about $1.35 / 1M tokens. Refer to NVIDIA’s official page for the exact rate.
- What context window does Nemotron 3 Ultra (batch) support? Context window is 512.3K, max output about 461.1K, suitable for long documents, multi-turn chat and complex reasoning.
- What scenarios is Nemotron 3 Ultra (batch) best for? Based on its capability and pricing, Nemotron 3 Ultra (batch) fits best: Deep Reasoning & Analysis. See the Use Cases section for details and cost estimates.
- What input/output modalities does Nemotron 3 Ultra (batch) support? Supports Text modality.
- Is Nemotron 3 Ultra (batch) a free model? Nemotron 3 Ultra (batch) is billed per token, not a free model.
- Which provider offers Nemotron 3 Ultra (batch)? Nemotron 3 Ultra (batch) is offered by NVIDIA.
More models from NVIDIA
Full info:Nemotron 3 Ultra (batch)