Brief IA

Google DeepMind: Revolutionizing World Models

🤖 Models & LLM·Tom Levy·

Google DeepMind: Revolutionizing World Models

Google DeepMind: Revolutionizing World Models
Key Takeaways
1Tom Zahavy from Google DeepMind claims that language models cannot initiate scientific revolutions.
2According to Zahavy, these models lack the cognitive mechanisms to generate radical innovations.
3World models are presented as a promising alternative for major scientific advancements.
💡Why it mattersWorld models could transform scientific research by overcoming the current limitations of language models.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Google Deepmind: The Revolution of World Models

Language models may not trigger scientific revolutions, but world models could. In a position paper titled "LLMs Cannot Leap," Tom Zahavy from Google Deepmind argues that they cannot. They lack the cognitive mechanism necessary to create something truly new.

Zahavy builds his argument on a framework outlined by Albert Einstein in a letter to his friend Maurice Solovine. Discovery, Einstein wrote, is a cycle: sensory experience leads to an intuitive "leap" towards axioms, and from there, logical deduction produces testable conclusions. Axioms are the unproven fundamental hypotheses of a theory.

AI Handles Two of the Three Types of Reasoning

To identify where the gap lies, Zahavy relies on a classic distinction made by philosopher Charles Sanders Peirce, who categorized all reasoning based on how it connects rules, cases, and outcomes.

  • Deduction derives guaranteed conclusions from fixed rules, such as executing a program that produces a provably correct output.

  • Induction identifies patterns in data: observing a thousand white swans and generalizing that all swans are white.

  • Abduction is the creative leap. It invents a cause to explain a surprising phenomenon.

This third form is where Zahavy sees the true bottleneck, and he draws a line between two levels. Ordinary abduction selects the most plausible explanation from a set of known candidates, like a doctor linking symptoms to a disease. Language models can do this, he concedes. The more challenging version is what he calls "manipulative abduction": inventing a cause for which no linguistic model yet exists. This, he argues, is the true bottleneck of scientific invention, and machines cannot do it.

Induction and deduction, according to the paper, are within reach. Language models already excel at recognizing statistical patterns and are quickly conquering formal derivation as well. Systems like AlphaProof, Gemini, and GPT-5 are now achieving gold-level scores on International Mathematical Olympiad problems. Zahavy even concedes that a language model could likely derive general relativity if given Einstein's hypotheses as a starting point. But formulating those hypotheses in the first place, making the manipulative leap to get there, remains the bottleneck.

Why Machines Struggle to Make This Leap

Zahavy illustrates why machines struggle with this leap using the same theory: AI models typically learn by comparing their predictions to reality and adjusting based on error, the gap between prediction and outcome. Without a detectable error, there is nothing for the system to work with. And this is the situation Einstein faced, Zahavy argues.

When Einstein was working, there was no data crisis. Newtonian physics had been confirmed with extreme precision. The only known anomaly, a slight shift in Mercury's orbit, had been attributed to a hypothetical hidden planet called "Vulcan." An optimization-focused AI would have had no reason to overturn physics, Zahavy contends. Following the logic of the argument, it would have done what astronomers of the time did: invent an additional planet to explain the small divergence, rather than rethinking space and time. The data confirming Einstein's theory, such as Eddington's measurement of light deflection, only arrived years after the theory was formulated.

A Leap Requires a Body

So, where does the manipulative abduction that led Einstein to his axioms come from? Zahavy points to Einstein's "happiest thought": the freely falling observer who no longer feels gravity. This intuition comes from embodied simulation, with Einstein mentally playing through a physical sensation rather than sticking to equations. He imagined a physicist inside an accelerating elevator in space and concluded that acceleration and gravity are indistinguishable from the inside.

Zahavy draws a parallel with Archimedes, who did not discover his principle of buoyancy through calculation, but, as the story goes, through the physical sensation of water rising when he entered a bathtub. In both cases, a fundamental principle emerged that did not yet exist in the language of the time.

Language models lack exactly this sensory grounding. Zahavy compares them to John Searle's thought experiment, the "Chinese Room," where a person mixes Chinese characters according to a manual without understanding a single word. Language models mix symbols of physics in the same way, without access to the physical experience that gives meaning to those symbols.

Sakana's AI Scientist and Deepmind's AlphaEvolve impressively automate scientific workflows. But the AI Scientist only recombines existing concepts, while AlphaEvolve brilliantly optimizes but needs a clear error signal that it can reduce step by step. Einstein never had that signal. Neither system, Zahavy argues, can make the leap to an entirely new framework of thought.

World Models as a Path to Abduction

As a possible path, Zahavy points to physically coherent world models. He draws a line here: video generators like Veo simply predict the next most probable frame. An apple falling falls not because the model understands gravity, but because falling is the most common continuation in the training data. It is still just a pattern matching.

Controllable action world models like Genie, on the other hand, allow an agent to actively intervene in a simulation and perform counterfactual experiments, like mentally cutting an elevator cable. A "synthetic laboratory" like this could provide the necessary feedback loop to invent new axioms where no linguistic model yet exists.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.