Inkling
The first general-purpose model from Mira Murati's Thinking Machines Lab, released July 15, 2026 under Apache 2.0 with open weights on Hugging Face. A 975B-parameter mixture-of-experts that activates about 41B parameters per task, trained on 45 trillion tokens and reasoning natively over text, images and audio (text out). Exposes a controllable thinking-effort setting. Rather than claiming the capability crown, Thinking Machines positions Inkling as a customizable foundation for teams that want control over model behaviour; it beats NVIDIA's Nemotron 3 Ultra on several evaluations. There is no first-party per-token price. The rates shown are representative third-party inference pricing (OpenRouter, ≈$1.00/$4.05 per 1M); it is also served by Together AI, Fireworks, Modal, Databricks and Baseten, and can be self-hosted or fine-tuned through Thinking Machines' Tinker service.
Inkling strengths
- Apache 2.0 open weights, no usage restrictions
- 975B total / ~41B active parameters keeps inference tractable
- Controllable thinking-effort setting
- Native text, image and audio reasoning
- Built for fine-tuning and behavioural control, not just prompting
Pricing & context
| Context window | Up to 1M tokens |
| Input price /1M | ≈$1.00 |
| Output price /1M | ≈$4.05 |
| Modalities | text, image, audio |
Cost guide: a typical call of about 10K input + 2K output tokens costs roughly $0.018 at list prices. Worth modelling against cheaper tiers before committing high-volume traffic.
When to choose Inkling
Inkling is best for research teams and companies that need to fine-tune and fully control a large multimodal model on their own infrastructure. If your workload is more cost-sensitive, weigh it against gpt-oss-120b (≈$0.03 input /1M) first.
Inkling FAQ
How much does Inkling cost?
Inkling is priced at ≈$1.00 per 1M input tokens and ≈$4.05 per 1M output tokens (public API list price), with a Up to 1M tokens context window. A typical call of about 10K input and 2K output tokens costs roughly $0.018.
What is Inkling best for?
Inkling by Thinking Machines Lab is best for research teams and companies that need to fine-tune and fully control a large multimodal model on their own infrastructure.
How does Inkling pricing compare to Laguna M.1?
Inkling input costs ≈$1.00 per 1M tokens versus $0.20 for Laguna M.1, roughly 5.0x more expensive on input. Output is ≈$4.05 vs $0.40.
Is Inkling multimodal?
Inkling supports text, image, audio.
Other models
All models →| 01 | Claude Fable 5 | Anthropic | $10.00 | → |
| 02 | Claude Opus 5 | Anthropic | $5.00 | → |
| 03 | GPT-5.6 Sol | OpenAI | $5.00 | → |
| 04 | GPT-5.5 | OpenAI | $5.00 | → |