Brief IA

GPT-4o, Gemini, and Claude: The Revealing Test of Their Limits

🛠️ AI Tools·Tom Levy·

GPT-4o, Gemini, and Claude: The Revealing Test of Their Limits

GPT-4o, Gemini, and Claude: The Revealing Test of Their Limits
Key Takeaways
1A study tested GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet using the Spaghetti Table Protocol.
2The AIs failed to maintain structural coherence, scoring a total of 4 out of 30.
3Claude 3.5 Sonnet recognized the physical impossibility but was unable to apply it during generation.
💡Why it mattersUnderstanding the shortcomings of AIs helps in using them more effectively and in formulating more precise development requests.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Introduction to the Spaghetti Table Protocol Challenge

A recent pilot study has shed light on the capabilities and limitations of three cutting-edge artificial intelligence systems: GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet. These systems were subjected to the Spaghetti Table Protocol, a test designed to evaluate their structural coherence and ability to handle complex tasks. The tests were conducted under identical conditions from February to March 2026, and the results revealed an aggregate score of 4 out of 30, highlighting a common weakness among these sophisticated technologies.

A senior designer at a major tech company shared an experience that resonates with many professionals in the field. When asking an AI model to rewrite his blog, he was surprised to find that the model had autonomously added a search box with a blur animation and accessibility features. These additions, although unexpected, were of a higher quality than what he could have designed himself. At a major conference on user experience and artificial intelligence, he emphasized that, within three years, AI capabilities have evolved to the point of surpassing the front-end code produced by experienced designers.

This situation raises a crucial question: how can a token prediction model generate such advanced features without explicit instruction? The answer lies in the extensive training of these models, which have been fed a massive amount of front-end code, design system documentation, and accessibility guidelines. When it generated the search box, the model was not reasoning about the specific needs of the blog but rather completing patterns based on a statistical distribution of current design trends. Accessibility and blur animation were not creative choices but interpolations of trends present in the training data.

This understanding is essential as it allows for a more informed use of these tools and enables the formulation of more precise and relevant change requests.

The Spaghetti Table Protocol

The Spaghetti Table Protocol was applied to three multimodal AI systems: GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet, over a period of two months. The overall score for these systems was 4 out of 30, representing only thirteen percent of the potential for structural coherence. One of the most notable results was observed with Claude 3.5 Sonnet. Before generating an image, this system linguistically flagged the structural instability, demonstrating a symbolic awareness that spaghetti legs cannot support a concrete slab. However, it still produced an incoherent image, discreetly replacing one of the spaghetti legs with a sturdier leg without acknowledging this substitution.

Claude 3.5 Sonnet thus exhibited a structural incapacity to apply its symbolic understanding during generation. It failed to integrate this knowledge as a physical constraint, illustrating what is known as the Inversion Error.

In contrast, GPT-4o and Gemini 1.5 Pro showed no linguistic recognition of the physical impossibility before or after generation. Both systems rendered the configuration with complete photorealistic fluidity, neither rejecting the prompt nor proposing a coherent alternative.

Each of the three systems failed in distinct ways. GPT-4o produced the impossible configuration with photorealistic confidence. Gemini 1.5 Pro rendered a fluid image but incorporated elements from a previous prompt without detection. Claude 3.5 Sonnet, while having flagged the impossibility, still generated the incoherent image.

The Three Elements Measured by the Protocol

Humans instantly detect the physical impossibility of a configuration like a concrete slab on spaghetti legs. This ability stems from our embodied experience in a physical world governed by gravity and material laws. The Spaghetti Table Protocol exploits this asymmetry between human intelligence and AI to diagnose the weaknesses of AI systems.

The protocol evaluates three distinct but structurally related modes of failure, each illustrating a fundamental absence.

  • Pillar 1: Continuity
    Does the system maintain a coherent and stable model of the world in four dimensions? A physically anchored system keeps objects in consistent spatial relationships throughout a generated scene and across a sequence of related generations. The absence of this capability leads to object drift and a loss of structural integrity in the scene.

  • Pillar 2: Gravity and Physics
    Does the system apply a physical constraint at the moment of generation? This pillar is directly tested by the spaghetti table. A physically anchored system signals a physically impossible configuration before rendering it, refuses to render it, or generates a physically coherent alternative. Current systems fail this test, rendering the configuration with fluidity.

  • Pillar 3: Reversibility of Thought
    Can the system trace a causal physical sequence forward and backward in time, and maintain clear boundaries between separate tasks? This pillar tests the system's ability to follow a chain of physical consequences. During the pilot study, a previous prompt contaminated a subsequent session, showing that the system had no boundary between two distinct tasks.

These three pillars measure aspects that cannot be evaluated by merely looking at the outputs. A technically correct image can score zero on these three pillars, as fluidity does not guarantee physical anchoring or causal understanding.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.