Stratego: Ataraxos Surpasses Humans with Less Than $8,000

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Ataraxos, developed by researchers from Carnegie Mellon, NYU, Stanford, and MIT, defeated Pim Niemeijer in 20 official matches. The training cost less than $8,000, and a first superhuman performance in Stratego has been claimed. The same method extends to other games with incomplete information, while acknowledging a limit on scaling. The code has been released.
A claimed scope beyond the game, with an assumed limit
The abundance of hidden information is no longer a barrier for reinforcement learning and research. They add that this advancement could translate into various strategic decisions, provided there are fast and faithful simulators, citing financial markets, military conflicts, and negotiations as examples. However, the researchers specify a limit: their research only simulates a single learning step, which prevents indefinite performance improvement through additional computation, although they believe that more elaborate research could mitigate this point. The Ataraxos code is made public. Beyond Stratego, the same method has been applied elsewhere: new records in Hanabi, with two orders of magnitude less computation for the two-player version, success against the benchmarks PerfectDou and DouZero in Dou dizhu, and four series of 50 winning games in Stratego Barrage against top-ranked players.
A duel organized over three weeks and a human advantage in adaptation
The confrontation with Pim Niemeijer spanned three weeks to limit fatigue and give him time to prepare. Ataraxos won 15 matches, conceded 1 defeat, and achieved 4 draws. His opponent, a four-time world champion, fifteen-time Dutch national champion, two-time online world champion, and over 600 weeks ranked number one, is described by George Franka as the most accomplished player of the game. Niemeijer knew that Ataraxos would not adapt to his style and earned $1,000, plus bonuses for each win and draw. The effective success rate of 85% is described as unprecedented at the highest level, while Max Roelofs reminds us of the narrow margins at the top. The asymmetry of adaptation favored the human, and Vincent de Boer sees this as a major handicap for the AI. On the sidelines, during an exhibition at the 2025 World Championship, Ataraxos won 38 out of 40 games against tournament players. Several participants described a style that was difficult to read, punctuated by bluffs deemed too risky, poorly perceived stalling phases, and a recurring sense of "luck" related to piece placement. The matches against Niemeijer are available online.
Regularization and belief network guide every decision
Ataraxos learns without human data, through self-play, and introduces a regularization constraint to diversify placements and moves. The necessity of randomization in Stratego to avoid predictability is emphasized. The regularization pressure decreases over the course of training, with adapted learning steps: broad exploration followed by fine adjustments. The researchers explain that this dynamic avoids cycles and chaotic phenomena typical of games with incomplete information, comparing regularization to an energy reserve that, if depleted too early, leads to quick but fragile gains. With each move, a belief network anticipates the opponent's pieces, generates hypotheses of states, simulates options, and triggers localized learning specific to the decision. Previous work had abandoned this type of research deemed too complex given the amount of hidden information.
Limited resources and an impossible comparison with DeepNash
The training took place on 16 H100 GPUs over one week, supplemented by four days on four GPUs for the belief network. The researchers estimate the cost at less than $8,000 at 2025 prices and assess their needs at about 1/500 of the computation, 1/30 of the self-play games, and 1/100 of the training examples compared to DeepNash. They attribute these discrepancies to a dedicated GPU simulator and better sampling efficiency, in a framework of academic resources highlighted by Samuel Sokota. DeepNash, for its part, ran on 1,024 TPU nodes for two to three months, and the cost of Ataraxos is estimated between $3 and $4.5 million in 2025. A direct comparison was not possible: the researchers claim to have provided the infrastructure, but DeepMind responded that the DeepNash code was no longer usable. At the 2023 World Championship, DeepNash won 19 out of 28 games but lost to most of the top players, including Niemeijer.
Why Stratego endures: 40 hidden pieces and 10^33 states
Each player lines up 40 concealed pieces, and the rank is only revealed upon confrontation, with victory going to the one who captures the opponent's flag. The researchers estimate the number of configurations to be beyond 10^33 and remind us, through Samuel Sokota, that the value of an action also depends on history in games with incomplete information. It is in this context that Ataraxos, developed by an academic team, claims a first superhuman performance in Stratego and a clear victory against Pim Niemeijer in 20 official matches.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.