Kimi K3: Moonshot AI Challenges the West with a Competing Model

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Kimi K3: Moonshot AI Challenges the West with a Competing Model
Moonshot AI's Kimi K3 model, which is reportedly close to rivaling the best Western models, raises new doubts about the effectiveness of U.S. export controls. Even a strategist from OpenAI is impressed.
A week ago, the research firm SemiAnalysis stated that Chinese labs were "simply too poor in computing to truly reach the frontier," a claim highlighted by Deepmind employee Anika Somaia. Just days later, Moonshot AI, a startup with around 300 employees, launched Kimi K3, which, according to early assessments, is comparable to Anthropic's Opus 4.8 but still falls short of cutting-edge models like Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol. The extent of this gap remains unclear.
Somaia argues that the Western consensus, from export controls to the race for hundreds of billions in investment from hyperscalers, rests on a single assumption: that computing power determines capability.
However, scarcity has forced innovation. Moonshot AI's Mooncake stack for AI training was developed precisely because the startup did not have enough GPUs, Somaia explains. "A small lab with taste can compress the computing needed to create a top-tier model, even if it can't afford to run one."
Dylan Patel, founder of the hardware analysis firm SemiAnalysis, agrees. "What they have achieved with a small, extremely talented team, strong research in reinforcement learning, architecture, and data compensates for much of the computing deficit," he writes. But Patel also emphasizes that Chinese companies can easily rent GPUs outside of China, making some export restrictions ineffective.
A Google Deepmind Researcher Calls Kimi K3 "Terribly Good"
Western AI labs often accuse Chinese companies of data theft through distillation, where a smaller AI model learns from the outputs of a larger one, thus benefiting from the work of others and threatening the business models of Western AI labs. So far, distillation has been the preferred explanation for understanding how Chinese labs remain competitive despite reduced computing resources.
For Kimi K3, this explanation does not seem to hold. "These results seem impossible to explain solely by distillation," writes Michiel Bakker, an AI researcher at MIT and Google Deepmind, calling the model "terribly good." Google's flagship model, Gemini 3.5 Pro, has reportedly been delayed for months according to Bloomberg, as it fails to meet performance targets, particularly in coding, its main use case. The company's AI strategy is once again under scrutiny, and Google faces regulatory headwinds in AI research, particularly in Germany.
Dean W. Ball, head of strategic futures at OpenAI and a former government advisor, describes Kimi as an "excellent model" that, in agent-based coding sessions, would match "the best public models of Q1 2026." However, he also notes that it appears "very token-hungry," making it "not obvious to me that this model is actually cheap to run."
He is not wrong. According to Artificial Analysis, Kimi K3 costs an average of $0.94 per task. This is close to GPT 5.6 Sol at $1.04, but about half the cost of Opus 4.8 at $1.80. It remains cheaper than the best Western models, but the gap has narrowed compared to the previous version, and it is much more expensive than older Chinese open-weight models.
An OpenAI Strategist Warns of "Total AI Communism"
However, Ball expresses surprise that the Chinese government allows the release of such powerful models as open-source. He attributes 75% of this to strategic blindness, claiming that the CCP is "very much in the style of Yann LeCun" in its assessment of AI-related risks and does not see existential threats.
The rest is due to a lack of computing capacity for client-side inference, which makes the open-weight strategy an unintended byproduct of U.S. export controls. Companies also know that almost no one would pay for Chinese models below the threshold, Ball asserts.
Open-weight models are "inherently decelerationist," Ball argues, as they further slow investment in AI. A possible outcome of a world dominated by them would be "total AI communism," with AI as a public good provided by the state as digital infrastructure. This is what China proposes, according to Ball, who describes this scenario as a "dystopian landscape."
The fact that an OpenAI strategist criticizes open-weight models so harshly is not without personal interest. His company relies on a closed business model and faces increasing price pressure from providers like Moonshot AI and Deepseek.
Regulatory Fog Rather Than Direct Bans
Ball predicts that the Trump administration will create a regulatory risk around the use of Chinese open-weight models. There is no need to ban open source, he argues, calling it "one of the most foolish reasons in the discussion about AI policy." Authorities would only need to create enough uncertainty through "soft laws," such as having the Federal Reserve issue warnings about potential backdoors in Chinese AI models. The reasoning wouldn't even need to be well-founded.
The goal is a middle ground with enough risk to deter regulated companies from using Chinese models, without scaring hyperscalers to the point that startups migrate to less reputable providers. Ball expects the government to deploy a version of this strategy.
More Efficient AI Could Still Mean Increased Demand for Computing
Kimi's advancements do not necessarily mean that less computing power is required, but if that were the case, the enormous infrastructures of U.S. tech companies would seem unnecessary, which could trigger a stock market crash. The opposite is, however, more likely. The Jevons Paradox suggests that more efficient models lead to increased deployment of AI, which could actually generate even more demand for computing power.
According to SemiAnalysis, Kimi K3 has 2.8 trillion parameters and is so large that it does not fit on a single Nvidia DGX B200, even with FP4 quantization. It requires more powerful systems like the GB300 NVL72 or B300, each with 288 GB of memory per GPU.
Once again, the parallels with Deepseek are hard to ignore. At the time, skeptics predicted a surplus of computing and briefly shook the markets. Instead, the demand for computing power increased as reasoning models scaled up, ironically fueled in part by Deepseek's own models. As Google Deepmind CEO Demis Hassabis puts it, "No one in the world knows what will happen next."
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.