Inkling (batch)

Inkling · Batch tier (Batch)

Asynchronous batch tier: discounted offline batch requests with high throughput and relaxed latency — ideal for large-scale non-real-time workloads.

Best for: Large-scale offline inference, cost-sensitive batch jobs, backfills

This page covers the Inkling Batch tier, sharing the underlying model capability with its real-time standard tier; the main difference is in billing mode and invocation method. See the Inkling standard real-time tier.

Provider:Thinking Machines

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Key specs

Token pricing

Modalities

Text, Image, Audio

Use Cases

When to pick this model

Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.

FAQ

More models from Thinking Machines

Full info:Inkling (batch)