Brief IA

Murati's Inkling: An Ambitious Model Against China

🎨 Creative AI·Tom Levy·

Murati's Inkling: An Ambitious Model Against China

Murati's Inkling: An Ambitious Model Against China
Key Takeaways
1Mira Murati, former CTO of OpenAI, founded Thinking Machines Lab and launched Inkling.
2Inkling, with 975 billion parameters, surpasses American models but not Chinese ones.
3The model is offered starting at $1.87 per million tokens, targeting fine-tuning.
💡Why it mattersInkling highlights the intense competition in AI, where even significant advancements can be overshadowed by international competitors.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Inkling by Murati: An Ambitious Model Facing China

Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, has launched Inkling, a multimodal open-weight model featuring 975 billion parameters that natively processes text, images, and audio.

According to the analysis platform Artificial Analysis, Inkling is currently the most powerful open-weight model in the United States, surpassing competitors like Kimi K2.6 and DeepSeek v4 Flash max on agentic tasks while demonstrating high efficiency in terms of tokens.

Despite its solid performance on benchmarks, the model exhibits notable weaknesses in factual accuracy, with a hallucination rate of 63%, and its cost is higher than that of comparable Chinese models.

Thinking Machines Lab has shipped its first production-ready language model. Inkling is a Mixture-of-Experts Transformer with a total of 975 billion parameters, of which 41 billion are active at any given time. It is the first model from the startup founded by Mira Murati, who played a key role in the development of ChatGPT.

Fine-Tuning as a Business Model

Unlike many other open-source AI models, Inkling natively handles text, images, and audio and supports a context window of up to one million tokens. The weights are available for free on Hugging Face. Thinking Machines also offers access via Tinker, its platform for adapting AI models to specific tasks.

The company positions Inkling as a flexible base model for customization. "Inkling is not the most powerful model available today," the announcement states. Thinking Machines expects that the combination of multimodal support, efficient processing, and fine-tuning options will set the model apart.

Thinking Machines claims to have pre-trained Inkling on 45 trillion tokens of text, images, audio recordings, and videos, whether public or synthetic. The training dataset also includes public data that "may be subject to intellectual property protection." The company used the Chinese AI model Kimi K2.5, among other methods, to generate synthetic data. Kimi K2.5 also served as a basis for the coding model from Cursor. More technical details are available in the model's specification sheet.

Inkling Leads American Open Models but Lags Behind Top Chinese Models

According to the benchmarking platform Artificial Analysis, Inkling debuts with a score of 41 on the Artificial Analysis Intelligence Index. This makes it the highest-performing open-weight model from an American lab. It ranks three points above the previous leader, Nemotron 3 Ultra at 38, and well ahead of Gemma 4 31B at 29 and gpt-oss-120b at 24.

Inkling achieves a score of 41 on the Artificial Analysis Intelligence Index, making it the highest-performing American open-weight model.

On GDPval-AA v2, an agent-based benchmark simulating cognitive work tasks, Inkling reaches an Elo ranking of 1,238. It outperforms Kimi K2.6 at 1,190 and DeepSeek v4 Flash max at 1,189. Inkling also scores 24% on the banking benchmark Tau-3, ahead of Kimi K2.6 at 21% and DeepSeek v4 Flash max at 23%.

Inkling surpasses Kimi K2.6 and DeepSeek v4 Flash max on agent-based cognitive work tasks.

Inkling shows rather mediocre performance in terms of factual accuracy. Artificial Analysis assigns the model a score of only +2 on its AA Omniscience benchmark. This places it below the highest-performing open-weight models, although it remains above other American models such as Nemotron 3 Ultra at -1. Inkling's accuracy is 40%, while its hallucination rate is 63%. These results may limit its use in applications requiring highly accurate information.

Inkling scores +2 on AA Omniscience, with an accuracy of 40% and a hallucination rate of 63%.

With a context window of 64K, Inkling costs $1.87 per million input tokens and $4.68 per million output tokens. This is slightly more than comparable open-source Chinese models like GLM-5.2 and DeepSeek v4, which offer similar or better performance on text and coding tasks. For context windows of up to 256,000 tokens, prices rise to $3.74 for input, $0.748 for cached input, and $9.36 for output.

However, Inkling uses fewer output tokens than comparable open-weight models. According to Artificial Analysis, it uses an average of 25,000 output tokens per task on the Intelligence Index. GLM-5.2 max uses 43,000, Kimi K2.6 uses about 38,000, and DeepSeek v4 Pro max uses about 37,000 tokens on the same tasks.

Thinking Machines claims that Inkling offers a continuously adjustable "thinking effort." Users can choose their preferred balance between cost and performance, reducing token usage while maintaining the same quality of output.

Inkling-Small Outperforms the Larger Model on Some Benchmarks

Thinking Machines also presents Inkling-Small, a more compact model with 276 billion parameters in total and 12 billion active parameters. The smaller model delivers similar or better results than Inkling on several benchmarks.

Inkling-Small scores 88.3% on GPQA Diamond, compared to 87.2% for Inkling. On the HLE benchmark with tools, it scores 46.6%, slightly ahead of Inkling at 46.0%. Thinking Machines attributes these results to changes in the pre-training data and training process. The company plans to release the full weights once testing is complete.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.