Inkling Small (batch)

Inkling Small · Batch tier (Batch)

Asynchronous batch tier: discounted offline batch requests with high throughput and relaxed latency — ideal for large-scale non-real-time workloads.

Best for: Large-scale offline inference, cost-sensitive batch jobs, backfills

This page covers the Inkling Small Batch tier, sharing the underlying model capability with its real-time standard tier; the main difference is in billing mode and invocation method. See the Inkling Small standard real-time tier.

Provider:Thinking Machines

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Key specs

Token pricing

Modalities

Text, Image, Audio

Use Cases

When to pick this model

Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.

FAQ

More models from Thinking Machines

Full info:Inkling Small (batch)