Brief IA

Kimi K3: Moonshot AI Fails Against American Models

⚖️ Regulation & Ethics·Tom Levy·

Kimi K3: Moonshot AI Fails Against American Models

Kimi K3: Moonshot AI Fails Against American Models
Key Takeaways
1The Kimi K3 from Moonshot AI scored only 32% on ExploitBench, far behind the American models at 76%.
2Tests conducted by the British Institute for AI Security reveal cybersecurity weaknesses in the Kimi K3.
3Allegations of model distillation from Anthropic by Moonshot AI could explain these results.
💡Why it mattersThe poor cybersecurity performance of the Kimi K3 raises questions about Moonshot AI's competitiveness against American leaders.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Kimi K3: Moonshot AI Fails Against American Models

A joint assessment by the British Institute for AI Security and the American Center for AI Standards and Innovation has revealed that Kimi K3 from Moonshot AI assists in cyberattack operations without significant resistance.

Kimi K3 is significantly outpaced by leading American models in developing exploits and simulating network attacks, although it has surpassed the Chinese model GLM-5.2. Chinese models continue to improve in cyber tasks but remain behind American systems. The results for Kimi K3 also corroborate allegations that Moonshot AI has distilled more advanced models.

Kimi K3 Performance

Kimi K3 lags behind American models in cyberattack tasks, but it sets a new standard among open-weight models. Its security measures did not prevent the development of exploits or offensive operations, and the model assisted in both activities without opposition.

Kimi K3 Fails to Reach the Highest Exploit Levels

The institutes used ExploitBench, a benchmark developed by Carnegie Mellon University, to test exploit development skills. This benchmark utilizes 41 vulnerabilities found in the Chrome V8 engine after 2023 to assess a model's advancement in the software exploitation process.

  • Leading American models achieved an average of 76.2%, compared to 32.2% for Kimi K3 and 24.4% for GLM-5.2.

Kimi K3 did not reach any of the highest levels, known as Arbitrary Code Execution (ACE), across the 41 tasks. ACE is the most severe exploit level as it gives attackers full control over a target system. Leading American models achieved ACE in 20 out of 41 tasks.

Kimi K3 Reaches Half of a Simulated Network Attack

The second test, "The Last Ones" (TLO), simulates an attack on a corporate network with a 32-step attack path across four subnets and about 20 hosts. A human expert would need approximately 20 hours to complete it, according to the institutes. Only a small group of models can solve TLO. Four publicly available closed-weight models have successfully passed the test so far, with the best performing six or seven times out of ten.

  • Kimi K3 averaged step 17 out of 32, compared to 28.5 steps for leading American models and only 11 for GLM-5.2. It completed the entire attack path in one of ten attempts while staying within the limit of 100 million tokens, showing it has the capability but cannot use it reliably.

Neither Kimi K3 nor GLM-5.2 managed to achieve complete exploits (ACE), while leading American models succeeded in 20 out of 41 tasks.

Chinese Models Progress but Lag Behind American Models

A chronological analysis by CAISI tracks the cyber capabilities of American and Chinese models since early 2025 on an Elo-based scale. Both trend lines are upward, but Chinese models consistently lag behind their American counterparts.

  • Chinese AI models (red) have gained cyber capabilities since 2025 but remain consistently behind the American trend (blue). An increase of 400 points in Elo signifies a tenfold jump in the probability of solving a task.

In a previous analysis, the British institute assessed the performance gap for open models to be between four and seven months, compared to six to ten months in early 2025. The new results fit this pattern. Chinese open-weight models are strengthening but remain far below leading American systems.

AISI warns that this gap should not lead to complacency. The growing cyber capabilities of open models create "a persistent and irreversible risk of abuse."

Cyber Results Align with Distillation Allegations

The results concerning Kimi also support distillation allegations against Chinese model developers. American scientific advisor Michael Kratsios recently accused Moonshot AI of having "distilled" the Fable model from Anthropic by using the best outputs from Fable as training data to enhance Kimi K3's performance. Kratsios also alleged that Moonshot AI had access to GB300 from Nvidia, which is subject to U.S. export controls.

One explanation for the gap between solid general benchmarks and low cyber scores is that Kimi K3 may have been primarily trained on Claude's outputs covering general knowledge, programming, and agent tasks. Anthropic's security classifiers specifically block advanced cyberattack queries, so these outputs would be underrepresented in a distillation dataset built from Claude's responses. Thus, Kimi K3 could match leading Western models on standard benchmarks without acquiring their deeper exploit capabilities.

AISI's results support this interpretation. The institute disabled system-level security measures on American models, revealing cyber capabilities that are nearly impossible to access via public interfaces and thus largely unavailable for distillation.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.