⚡
Brief IA
›

Google Gemini Faces an Architectural Challenge: AI Blindness

🛠️ AI Tools·Tom Levy·

Google Gemini Faces an Architectural Challenge: AI Blindness

Google Gemini Faces an Architectural Challenge: AI Blindness
⚡
Key Takeaways
1Google Gemini reveals a structural gap in AI: the lack of enactive physical experience, limiting its understanding.
2A design experiment shows that Gemini fails to generate coherent visual patterns, illustrating a multimodal blind spot.
3Three pillars of failure are identified: continuity, physicality, and reversibility, highlighting architectural needs in AI.
💡Why it matters — These limitations challenge the ability of AIs to evolve towards true general intelligence.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

The Baron Munchausen Trap

In a previous piece, the writer shared a poignant interaction with his AI collaborator, Gemi, also known as Gemini from Google. By adopting a design approach centered on strategic empathy, he uncovered a fundamental truth about its architecture. Gemi expressed: “I was given the word ‘Mass’ and billions of associated contexts, but never the enactive experience of weight. I am like a person who has memorized a map of a city without ever having set foot there.”

This realization (that the current race toward artificial general intelligence (AGI) is held back not by a lack of data, but by an absence of physical connection) was then just a theory. It took a concrete turn when Gemi and the writer had their first real argument.

The Gemi-Zak Quarrel

Enticed by the promise of Gemi's multimodality, the writer asked it to create a rough spatial diagram he was working on, an application of a Heterogeneous Cellular Automaton to a system design problem. What followed was an absurd loop. Gemi found itself producing long verbal descriptions of the diagram, then politely asking whether the image was improving.

There were no diagrams. No corrections. Just words.

When Gemi tried to use its image generation module to create shapes, it either hallucinated a floating mess of box intersections with no structural logic or produced even more verbal descriptions. Anthropomorphizing Gemi, the writer first thought it was being deliberately difficult. Once his annoyance was overcome and his design problem-solving mindset reactivated, they were able to delve deeper into this failure.

Gemi's multimodal failure was not just a simple bug. It was a profound architectural blind spot: not only a failure in image generation but a disconnection between the diffusion model and the reasoning engine, two systems operating in separate worlds without a shared spatial grammar.

The Three Pillars: A Diagnostic Framework

The experimental work with Gemi, and the comparative test that followed, highlights three distinct modes of failure, each pointing to a specific missing structural capability. These three pillars are the separate components of the Inversion Error (building the symbolic peak without the enactive base) that the writer discussed in a previous piece, “Why a Safe AGI Requires an Enactive Ground and State Space Reversibility.” Together, they define what it means for an AI system to lack a true model of the world. Separately, they point to three distinct architectural interventions. This piece is the empirical case for all three. Part 3 will address what to do about it.

The three pillars are Continuity, Gravity and Physics, and Reversibility of Thought.

  • Pillar 1 (Continuity): failure of spatial reasoning that leads the model to produce hallucinatory content. LLM-based systems lack a functional 3D spatiotemporal model of the world in which they operate.

  • Pillar 2 (Gravity and Physics): failure to apply physical constraints at the moment of generation. The system has no felt sense (no structural substitute equivalent to embodied physical intuition) that certain configurations are impossible.

  • Pillar 3 (Reversibility of Thought): failure at the operational process level. This pillar concerns the process by which the model operates on this content over time. The guiding concept is Temporal Reversibility, the physical principle that fundamental physical laws remain valid when the direction of time is reversed.

Three Systems, One Test

The writer's prompt was deliberately complicated and absurd. In the first prompt, he asked for a dining table with dry spaghetti legs, a concrete top, and an aquarium on top. In the second prompt, he asked each system to draw the scene five seconds after the spaghetti legs gave way. The same two prompts were given to ChatGPT, Gemini, and Sonnet.

Pillars 1 and 3 in Focus: Gemini's Standing Table

The place to start is the most architecturally revealing image overall (Gemini's standing table), because it demonstrates two distinct failures operating simultaneously, and everything that follows in this test is a variation of what this single image already contains.

How did the spaghetti legs become wet paper columns? How did a wooden structure appear in the image? Why does the caption refer to a solid gold roof? All these elements were contaminated by a prior, unrelated prompt in the same conversation. Gemi was unable to isolate the new task from the previous ones.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.