Inkling

Thinking Machines Labtextimageaudio

The first general-purpose model from Mira Murati's Thinking Machines Lab, released July 15, 2026 under Apache 2.0 with open weights on Hugging Face. A 975B-parameter mixture-of-experts that activates about 41B parameters per task, trained on 45 trillion tokens and reasoning natively over text, images and audio (text out). Exposes a controllable thinking-effort setting. Rather than claiming the capability crown, Thinking Machines positions Inkling as a customizable foundation for teams that want control over model behaviour; it beats NVIDIA's Nemotron 3 Ultra on several evaluations. There is no first-party per-token price. The rates shown are representative third-party inference pricing (OpenRouter, ≈$1.00/$4.05 per 1M); it is also served by Together AI, Fireworks, Modal, Databricks and Baseten, and can be self-hosted or fine-tuned through Thinking Machines' Tinker service.

Inkling strengths

  • Apache 2.0 open weights, no usage restrictions
  • 975B total / ~41B active parameters keeps inference tractable
  • Controllable thinking-effort setting
  • Native text, image and audio reasoning
  • Built for fine-tuning and behavioural control, not just prompting

Pricing & context

Context windowUp to 1M tokens
Input price /1M≈$1.00
Output price /1M≈$4.05
Modalitiestext, image, audio

Cost guide: a typical call of about 10K input + 2K output tokens costs roughly $0.018 at list prices. Worth modelling against cheaper tiers before committing high-volume traffic.

When to choose Inkling

Inkling is best for research teams and companies that need to fine-tune and fully control a large multimodal model on their own infrastructure. If your workload is more cost-sensitive, weigh it against gpt-oss-120b (≈$0.03 input /1M) first.

Inkling FAQ

How much does Inkling cost?

Inkling is priced at ≈$1.00 per 1M input tokens and ≈$4.05 per 1M output tokens (public API list price), with a Up to 1M tokens context window. A typical call of about 10K input and 2K output tokens costs roughly $0.018.

What is Inkling best for?

Inkling by Thinking Machines Lab is best for research teams and companies that need to fine-tune and fully control a large multimodal model on their own infrastructure.

How does Inkling pricing compare to Laguna M.1?

Inkling input costs ≈$1.00 per 1M tokens versus $0.20 for Laguna M.1, roughly 5.0x more expensive on input. Output is ≈$4.05 vs $0.40.

Is Inkling multimodal?

Inkling supports text, image, audio.

Other models

All models →
01Claude Fable 5Anthropic$10.00
02Claude Opus 5Anthropic$5.00
03GPT-5.6 SolOpenAI$5.00
04GPT-5.5OpenAI$5.00