Nemotron 3 Ultra (batch)

Nemotron 3 Ultra · Batch tier (Batch)

Asynchronous batch tier: discounted offline batch requests with high throughput and relaxed latency — ideal for large-scale non-real-time workloads.

Best for: Large-scale offline inference, cost-sensitive batch jobs, backfills

This page covers the Nemotron 3 Ultra Batch tier, sharing the underlying model capability with its real-time standard tier; the main difference is in billing mode and invocation method. See the Nemotron 3 Ultra standard real-time tier.

Provider:NVIDIA

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Key specs

Token pricing

Modalities

Text

Use Cases

When to pick this model

Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.

FAQ

More models from NVIDIA

Full info:Nemotron 3 Ultra (batch)