Inkling (batch)
Inkling · Batch tier (Batch)
Asynchronous batch tier: discounted offline batch requests with high throughput and relaxed latency — ideal for large-scale non-real-time workloads.
Best for: Large-scale offline inference, cost-sensitive batch jobs, backfills
This page covers the Inkling Batch tier, sharing the underlying model capability with its real-time standard tier; the main difference is in billing mode and invocation method. See the Inkling standard real-time tier.
Provider:Thinking Machines
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Key specs
- Context:524.3K
- Max output:471.9K
- Tokenizer:Other
- Released:2026-07-17
Token pricing
- Input price:$1.00 / 1M tokens
- Output price:$4.05 / 1M tokens
- Cache read:$0.170 / 1M tokens
- Blended price:$1.76 / 1M tokens
Modalities
Text, Image, Audio
Use Cases
- Deep Reasoning & Analysis:Process long legal docs, academic papers, financial reports; multi-hop reasoning and evidence-chain tracing across huge contexts.(≈ $0.032/call)
- Vision Understanding & Generation:Recognize data trends in charts, OCR scanned docs, analyze photos and produce structured reports.(≈ $0.024/call)
- Voice Interaction:Speech recognition & transcription, mixed CN/EN and dialect support — meeting notes, QA.
- Chat & Support:FAQ auto-response, multi-turn dialogue, emotion detection — e-commerce, finance, gov.(≈ $0.0081/call)
When to pick this model
Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.
- Long max output (>=30K tokens): strong fit for deep report / essay / document-drafting workflows.
- Multimodal (text+image): fit for image captioning, OCR, chart QA, vision-grounded reasoning.
- Audio support: fit for speech-to-text / voice agent / transcript analysis.
FAQ
- What is the token pricing of Inkling (batch)? Input price $1.00 / 1M tokens, output price $4.05 / 1M tokens, blended about $1.76 / 1M tokens. Refer to Thinking Machines’s official page for the exact rate.
- What context window does Inkling (batch) support? Context window is 524.3K, max output about 471.9K, suitable for long documents, multi-turn chat and complex reasoning.
- What scenarios is Inkling (batch) best for? Based on its capability and pricing, Inkling (batch) fits best: Deep Reasoning & Analysis, Vision Understanding & Generation, Voice Interaction. See the Use Cases section for details and cost estimates.
- What input/output modalities does Inkling (batch) support? Supports Text, Image, Audio modality.
- Is Inkling (batch) a free model? Inkling (batch) is billed per token, not a free model.
- Which provider offers Inkling (batch)? Inkling (batch) is offered by Thinking Machines.
More models from Thinking Machines
Full info:Inkling (batch)