MiMo-V2.5
Provider:Xiaomi
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
Key specs
- Context:1.1M
- Max output:131.1K
- Tokenizer:Other
- Released:2026-04-22
Token pricing
- Input price:$0.140 / 1M tokens
- Output price:$0.280 / 1M tokens
- Cache read:$0.0028 / 1M tokens
- Blended price:$0.175 / 1M tokens
Modalities
Text, Audio, Image, Video
Use Cases
- Deep Reasoning & Analysis:Process long legal docs, academic papers, financial reports; multi-hop reasoning and evidence-chain tracing across huge contexts.(≈ $0.0022/call)
- Vision Understanding & Generation:Recognize data trends in charts, OCR scanned docs, analyze photos and produce structured reports.(≈ $0.0017/call)
- Voice Interaction:Speech recognition & transcription, mixed CN/EN and dialect support — meeting notes, QA.
When to pick this model
Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.
- Low input + output price: solid fit for high-frequency calls and batch processing.
- Low blended price: budget-friendly for fleet-scale deployment.
- Long context (>=1M tokens): well-suited for long-doc QA, code-base analysis, multi-turn agents.
- Long max output (>=30K tokens): strong fit for deep report / essay / document-drafting workflows.
- Multimodal (text+image): fit for image captioning, OCR, chart QA, vision-grounded reasoning.
- Audio support: fit for speech-to-text / voice agent / transcript analysis.
- Video input support: fit for video understanding, frame QA, content moderation.
FAQ
- What is the token pricing of MiMo-V2.5? Input price $0.140 / 1M tokens, output price $0.280 / 1M tokens, blended about $0.175 / 1M tokens. Refer to Xiaomi’s official page for the exact rate.
- What context window does MiMo-V2.5 support? Context window is 1.1M, max output about 131.1K, suitable for long documents, multi-turn chat and complex reasoning.
- What scenarios is MiMo-V2.5 best for? Based on its capability and pricing, MiMo-V2.5 fits best: Deep Reasoning & Analysis, Vision Understanding & Generation, Voice Interaction. See the Use Cases section for details and cost estimates.
- What input/output modalities does MiMo-V2.5 support? Supports Text, Audio, Image, Video modality.
- Is MiMo-V2.5 a free model? MiMo-V2.5 is billed per token, not a free model.
- Which provider offers MiMo-V2.5? MiMo-V2.5 is offered by Xiaomi.
More models from Xiaomi
Full info:MiMo-V2.5