AI Puzzles: Rapid Progress, Persistent Blind Spots

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Recent studies and concrete examples highlight the strengths and weaknesses of models. They excel at simple puzzles but stumble as the scale increases, struggle with spatial and visual grids, and fall prey to the familiarity of statements. Conversely, their performance has surged in certain recent games.
Beyond six elements, studies observe a drop-off
Several studies point to a link between the scale of a puzzle and the success of the models. Researchers from Apple report that LLMs perform well on simple versions of the Tower of Hanoi and river crossing puzzles, but their performance degrades when the configuration reaches six elements or more. The Tower of Hanoi requires moving disks one by one without placing a larger disk on a smaller one, and the river crossing puzzles require a group to adhere to specific rules. Another study conducted at the University of Washington, Stanford, and the Allen Institute for AI observes similar difficulties with logical grids, where one must deduce attributes of individuals from clues. Apple has widely disseminated its findings, while commentators have debated whether this is a limit inherent to LLM reasoning or ordinary errors due to complexity. Practical statements extend these findings, inviting planning for crossing routes or completing a uniquely solvable grid.
Familiarity of statements leads to memory errors
State-of-the-art LLMs have accumulated a vast volume of facts and excel in quiz-like contexts, but this memory can backfire. If a puzzle bears a strong resemblance to an example encountered during training, they may overlook important distinctions and produce a response based on their memory. In 2024, researchers from Google and the University of Illinois at Urbana-Champaign trained and evaluated models on very close variants of Knights and Knaves, a game where some characters are consistently truthful and others consistently liars. A similar operation is assumed for SimpleBench, whose questions evoke more sophisticated problems that were likely encountered during training. Where humans detect the trap, leading models stumble. The associated instructions recall the rules of the characters and encourage careful reading of the statements.
3D space resists language models
Spatial reasoning remains an area of advantage for humans. Mental rotation puzzles, familiar from IQ tests, remain difficult for models that are otherwise capable of analyzing visual inputs. Despite discussions around "world models," LLMs do not seem to manipulate three-dimensional objects as seasoned professionals do in spatial contexts. Proposed exercises require identifying the same object viewed from a different angle, with a single correct answer.
Grids and abstract rules: ARC-AGI exposes flaws
The difficulties are not limited to 3D: in two dimensions as well, puzzles pose problems. This is crucial on ARC-AGI, a benchmark of puzzles where one must infer general rules from examples. Models perform better when grids are provided as chains of numbers encoding colors rather than as images. Studies suggest that they then resort to Byzantine rules that are not easily generalizable, while humans mobilize simple visual concepts. Despite these handicaps, their scores have significantly improved over the past year, although some items continue to confuse them. A typical exercise involves deducing the transformation rule between pairs of grids before completing a fourth, regardless of orientation.
When human biases turn the trap
AI is not alone in making mistakes: cognitive limits inherent to humans have been exploited by psychologists in series of problems that reverse the advantage observed elsewhere. In these cases, human responses are often instinctive while models produce more deliberate answers. Some statements play on common errors in intuitive calculation, while others induce responses that collapse upon careful reading. The instructions call for a quick response, which exacerbates these reflexes.
From checkers to Connections: recent milestones and figures
Since Arthur Samuel's algorithm popularized "machine learning" in 1959 with checkers, games like chess or Go have served as testing grounds. Puzzles play the same role in assessing fine skills. On the New York Times' Connections, a team from Columbia observed that by the end of 2024, even the best models solved only 18%, while by early 2025, some were almost consistently successful. The puzzles reveal both progress and flaws: subtleties of statements, fragility in the face of visuals. The proposed exercises invite observers to note these contrasts, each puzzle illustrating a difference in cognition, while offering the reader the opportunity — for now — to surpass the models.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.