Brief IA

Kimi K3 Challenges GPT-5.6, Sol, and Fable 5 with Its Revolutionary Model

🤖 Models & LLM·Tom Levy·

Kimi K3 Challenges GPT-5.6, Sol, and Fable 5 with Its Revolutionary Model

Kimi K3 Challenges GPT-5.6, Sol, and Fable 5 with Its Revolutionary Model
Key Takeaways
1Kimi has unveiled K3, a multimodal open-weight model with 2.8 trillion parameters.
2K3 approaches the performance of Claude Fable 5 and GPT 5.6 Sol, surpassing Opus 4.8 and GLM 5.2.
3The K3 model is more expensive than its predecessors, with weight release scheduled for July 27.
💡Why it mattersK3 could redefine the standards of AI, marking a turning point in global technological competition.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Kimi K3 Challenges GPT-5.6 Sol and Fable 5 with Its Revolutionary Model

Kimi has launched K3, a multimodal open-weight model built on a mixture of experts architecture featuring 896 experts, 2.8 trillion parameters, and a context window of one million tokens. The full weights are expected by the end of July.

In Kimi's own benchmarks, K3 comes close to Claude Fable 5 and GPT 5.6 Sol, but significantly outperforms all other tested systems. Independent tests conducted by Artificial Analysis largely confirm these results, although K3's hallucination rate has increased compared to its predecessor.

At $3 per million input tokens and $15 per million output tokens, K3 is much more expensive than its predecessor but comparable to mid-range Western models like Sonnet 5. The costs per task amount to about $0.94, similar to GPT-5.6 Sol and roughly half the price of Opus 4.8.

Kimi describes K3 as the first open model in the range of about 3 trillion parameters. The model targets long-duration programming tasks, knowledge work, and complex reasoning.

In Kimi's benchmarks, K3 lags behind the leading proprietary models Claude Fable 5 and GPT 5.6 Sol, but surpasses all other tested systems, including Claude Opus models and the Chinese rival GLM-5.2. All results come from Kimi and were obtained at maximum or high reasoning intensity, according to the company.

Kimi K3 wins two of the six programming benchmarks and ranks second or third in the others.

Performance and Evaluations

Among general agents, Kimi K3 wins three of the six tests. Fable 5 takes the top spot in both visual agent tests.

Across all 35 tests, K3 has taken first place about seven times and ranked second or third in most others. Fable 5 has won the most individual tests. In nearly all benchmarks, K3 has significantly outperformed Opus 4.8, GPT 5.5, and GLM 5.2. Depending on the benchmark, one of three agent systems was used: KimiCode, Claude Code, or Codex, meaning that the results were not all collected under identical conditions.

Artificial Analysis has published its first evaluation of Kimi K3. The model scores 57 on the Artificial Analysis Intelligence Index, placing it on par with Opus 4.8 and GPT-5.5, but still behind Fable 5 and GPT-5.6 Sol. This largely aligns with Kimi's claims.

On agent tasks, K3 achieves an Elo ranking of 1,668 on GDPval v2, a significant jump from K2.6, which was 1,190. It surpasses GLM-5.2 (1,514), GPT-5.5 (1,494), and Claude Opus 4.8 (1,600), although it still falls short of Claude Fable 5 (1,760). K3 also ranks first on AutomationBench-AA, the Artificial Analysis version of the SaaS agent workflow evaluation from Zapier, with a score of 53%.

On AA-Briefcase, a private evaluation of long-term knowledge work, K3 achieves an overall Elo of 1,547, up 732 points from K2.6. Only Claude Fable 5 scores higher. Artificial Analysis describes K3 as well-balanced, with grid scores and analytical quality close to those of Fable 5. However, GPT-5.6 Sol remains ahead in terms of presentation quality.

K3's accuracy rate has improved, rising from 33% to 46% on the AA-Omniscience Index, increasing the overall score from +6 to +18. However, its hallucination rate has climbed from 39% to 51%, meaning K3 generates more responses even as it correctly answers more questions.

Use Cases and Architecture

According to Kimi, the primary use case for the model is long-duration software development with minimal human oversight. K3 is designed to analyze large codebases, coordinate terminal tools, and maintain focus on a task across many work steps.

The model combines programming with visual feedback: it examines screenshots, modifies code, and then checks the visible output. Kimi calls this closed-loop system "Vision in the Loop" and positions it as a foundation for game development, user interface design, and CAD.

K3 employs a mixture of experts architecture that activates only 16 of the 896 experts at a time. It is paired with a new attention architecture called Kimi Delta Attention, which, according to Kimi, allows for decoding up to 6.3 times faster for contexts of one million tokens. The "attention residues" would increase training efficiency by about 25% while adding less than 2% additional computational overhead.

Pricing and Availability

According to Kimi's API documentation, one million input tokens costs $0.30 with caching and $3.00 without. One million output tokens, including reasoning, costs $15.00. These prices apply regardless of context length. Caching occurs automatically, making long, unmodified prefixes particularly useful for agents and large codebases.

This places K3 well above the price level of its predecessor K2.6, which officially costs $0.16 per million tokens with caching, $0.95 without, and $4.00 for output. Chinese providers no longer offer their top models at very low prices either.

However, K3 is much cheaper than the best Western models and falls more into the upper mid-range. The new Sonnet 5 from Anthropic, for example, also costs $3 per million input tokens and $15 for output, but offers lower performance.

K3 also uses fewer tokens than its predecessor. It required about 132 million output tokens to complete the nine evaluations, compared to about 166 million for K2.6, a 21% reduction while achieving 13 points more. Due to the much higher token prices, K3 will likely cost even more per task than K2.6 in most cases.

K3 is already available via Kimi.com, the mobile app for iOS, Android, and HarmonyOS, the desktop client Kimi Work (version 3.1.0 and later), and Kimi Code. On OpenRouter, the model is listed under the identifier "moonshotai/kimi-k3," although it is currently served only by Moonshot. The open weights are expected by the end of July.

For businesses, Kimi offers a separate version with member management and the ability to separate personal and professional accounts. A planned platform called Kimi Hosted Agent will provide isolated environments and runtimes for long-duration tasks. Interested users can sign up for the waitlist now.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.