Mercury 2.5
Provider:Inception Labs
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Key specs
- Context:260K
- Max output:65.5K
- Tokenizer:Other
- Released:2026-09-08
Token pricing
- Input price:$0.040 / 1M tokens
- Output price:$0.150 / 1M tokens
- Cache read:$0.0040 / 1M tokens
- Blended price:$0.068 / 1M tokens
Modalities
Text
Use Cases
- Deep Reasoning & Analysis:Process long legal docs, academic papers, financial reports; multi-hop reasoning and evidence-chain tracing across huge contexts.(≈ $0.0012/call)
When to pick this model
Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.
- Low input + output price: solid fit for high-frequency calls and batch processing.
- Low blended price: budget-friendly for fleet-scale deployment.
- Long max output (>=30K tokens): strong fit for deep report / essay / document-drafting workflows.
FAQ
- What is the token pricing of Mercury 2.5? Input price $0.040 / 1M tokens, output price $0.150 / 1M tokens, blended about $0.068 / 1M tokens. Refer to Inception Labs’s official page for the exact rate.
- What context window does Mercury 2.5 support? Context window is 260K, max output about 65.5K, suitable for long documents, multi-turn chat and complex reasoning.
- What scenarios is Mercury 2.5 best for? Based on its capability and pricing, Mercury 2.5 fits best: Deep Reasoning & Analysis. See the Use Cases section for details and cost estimates.
- What input/output modalities does Mercury 2.5 support? Supports Text modality.
- Is Mercury 2.5 a free model? Mercury 2.5 is billed per token, not a free model.
- Which provider offers Mercury 2.5? Mercury 2.5 is offered by Inception Labs.
More models from Inception Labs
Full info:Mercury 2.5