Brief IA

AI: Success in Math, Failures in Timekeeping, a Paradox Persists

🛠️ AI Tools·Tom Levy·

AI: Success in Math, Failures in Timekeeping, a Paradox Persists

AI: Success in Math, Failures in Timekeeping, a Paradox Persists
Key Takeaways
1AI models are winning gold medals at math Olympiads, showcasing superhuman abilities in complex tasks.
2Despite these successes, AI struggles with simple tasks like reading clocks, achieving only a 50.6% success rate.
3The Stanford AI Index 2026 report highlights the uneven performance of AI, a phenomenon referred to as "jagged intelligence."
💡Why it mattersThis performance disparity complicates the integration of AI into everyday tasks, limiting its practical utility.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Unmatched Performance at the Mathematics Olympiad

The AI Index 2026 report from Stanford HAI, published on April 13, 2026, in its ninth edition, highlights the exceptional performance of artificial intelligence models in demanding contexts. In particular, Google’s Gemini Deep Think model won the gold medal at the 2025 International Mathematical Olympiad. This model solved five out of six problems using only natural language within a time frame of 4 hours and 30 minutes, surpassing the previous year's performance, where a silver medal was obtained after translating the problems into formal language and several days of computation.

Surprising Failures in Simple Tasks

Despite these successes, AI shows notable weaknesses in simpler tasks. For example, on ClockBench, a test designed to evaluate models' ability to read analog clocks, the GPT-5.4 High model achieved only a 50.6% success rate, far below the 90.1% achieved by humans. The errors made by the models are significant, with a median deviation of 1 to 3 hours, compared to just 3 minutes for a human.

"Jagged Intelligence": A Persistent Challenge

The phenomenon of "jagged intelligence," described in the Stanford report, illustrates this disparity in performance. AI models excel in certain complex tasks but fail in more basic actions. This inequality is also evident in the fields of robotics and science. For instance, while robotic systems achieve a 89.4% success rate in simulation on RLBench, the Robot Learning Collective team, winner of the BEHAVIOR Challenge 2025, completed only 12.4% of realistic household tasks.

In the scientific domain, several models outperform human chemists on ChemBench on average, but fall below 20% on astrophysics replication and 33% on Earth observation questions.

Limitations in Model Training

The difficulties faced by AI are not solely due to a lack of training data. A 2025 study mentioned in the report attempted to improve model performance on clock reading using 5,000 synthetic images. Although the models made progress on familiar clock designs, they failed to generalize on unusual designs. The issue lies in how models interpret visual cues, particularly the confusion between hour and minute hands.

For digital professionals considering automating tasks, this jagged boundary has a direct implication. A model may appear impressive in a carefully chosen demo but falter on a seemingly simpler task. The report provides a new illustration with OSWorld, a benchmark that tests AI agents on real computer tasks (Ubuntu, Windows, macOS): performance increased from about 12% to 66.3% success in one year with Claude Opus 4.5, just six points shy of the human average. But this also means that about one in three tasks is still missed, on actions that computer science students complete in two minutes. In this context, testing models on one's own use cases remains the only true indicator of their operational utility.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.