Design of Ataraxos
Ataraxos consists of two interdependent self-play reinforcement learning processes, realized by transformer networks for set-up selection and move selection; a belief network trained on self-play games of the final selection networks; and a test-time search procedure that composes the move and belief networks.
Interdependent self-play processes
The foundation of Ataraxos is its self-play reinforcement learning module. This module comprises two separate but interdependent self-play processes associated with the two phases of Stratego. The first phase, in which players privately determine starting positions for their pieces—called set-ups—is handled by one self-play process. The second phase, in which players alternate moving their pieces, is handled by another. These processes learn separately but in tandem: the set-up selection process determines the initial boards for the move selection process; and the move selection process determines game outcomes that both processes use for policy updates. We chose this decomposition rather than an end-to-end model because, although an end-to-end approach would avoid making the phase decomposition explicit, it would require a single model and training pipeline to accommodate two unrelated tasks. As discussed in the sections ‘Reinforcement learning with transformers’, ‘Self-play training data generation’ and ‘Dynamically damped self-play’, with full Stratego training specifications provided in Supplementary Information, set-up learning favoured a decoder-only architecture, Monte Carlo estimation of expected returns and advantages, and higher learning rates and regularization temperatures. Move learning, by contrast, favoured an encoder-only architecture, λ-based expected return and advantage estimators (in which λ controls the weighting of multi-step estimates), an annealed learning-rate schedule, and lower regularization temperatures.
Reinforcement learning with transformers
To represent the policies and value functions for the two self-play processes, Ataraxos uses two transformers32, which we call the set-up network and the move network. These networks are diagrammed in Extended Data Fig. 1. Important choices included: parameterizing the set-up network as a decoder-only transformer, which allowed training on entire set-ups with single forward–backward passes; parameterizing the policy over moves using a key–query matrix product33, which learned faster than less sophisticated parameterizations; using learned absolute positional embeddings34 in both networks; and sizing the move network to balance sample efficiency against iteration speed, between which we observed substantial trade-offs.
Self-play training data generation
To generate self-play games, Ataraxos selects its set-ups and moves by directly sampling from the associated networks.
For the moves played during these games, Ataraxos computes the estimates of expected cumulants35 and advantages needed for training updates using λ returns36,37 (with distinct λ values), and trains only on the moves with large estimated advantage magnitudes38. Filtering in this way reduced the overall wall-clock time per reinforcement learning iteration by a factor of about 2.5, while simultaneously—in a phenomenon meriting further investigation—actually increasing both sample efficiency (with respect to the number of environment queries) and asymptotic performance.
For the set-ups, Ataraxos uses Monte Carlo returns (that is, final outcomes of games played by the current policy) to estimate these quantities, and does not use filtering. The superiority of Monte Carlo returns for advantage estimation is unusual in reinforcement learning, although similar behaviour has also been observed in language model reasoning39.
Dynamically damped self-play
Ataraxos trains on on-policy or close to on-policy self-play data by dynamically damping its learning dynamics. This allows it to leverage standard policy optimization tools to regularize and control the size of its policy updates, obviating the need for more onerous techniques, such as trajectory importance reweighting and policy averaging, for handling imperfect information.
For regularization, Ataraxos incorporates additional terms into the losses of its networks, as detailed in ‘Stratego implementation details’ in Supplementary Information. For the set-up network, it uses a maximum entropy term40; for the move network, it uses a myopic reverse Kullback–Leibler penalty towards the policy that selects a movable piece uniformly at random and then selects a legal move for this piece uniformly at random. Ataraxos anneals the coefficients of these regularization terms according to distinct power laws over training12. We found that this regularization functioned somewhat analogously to an energy reserve—annealing too cautiously left the playing ability of the model underdeveloped, while annealing too aggressively produced rapid initial gains but also collapsed the entropy of the model, depleting its capacity to learn thereafter and often making it easy to exploit.
For update size control, Ataraxos uses four mechanisms, which we found offered complementary benefits: a reverse Kullback–Leibler penalty to the data collection policy, importance ratio clipping41, gradient norm clipping42 and the learning rate of Adam43. Ataraxos anneals the learning rate for its move network according to a power law over the course of training. We found that scheduling this learning rate was crucial both for rapid learning early in training and for preventing plateauing later in training.
Belief modelling
To facilitate search, Ataraxos trains a belief network on every position of trajectories sampled from its final self-play policy. Given the information available to the player at a position, this network is trained with teacher forcing to maximize the log-likelihood of the true types of the opponent’s hidden pieces. The belief network processes the known information using a transformer encoder (similar to the move network) and predicts the types of unknown pieces autoregressively in row-major order using a transformer decoder. The resulting architecture is diagrammed in Extended Data Fig. 4. Ataraxos applies dropout44 to the belief network during training to aid generalization to out-of-distribution positions reachable by opponents that play very differently from its set-up and move networks, such as humans.
Test-time search via update equivalence
Given move and belief networks, the search Ataraxos uses is straightforward—it simply performs an additional damped self-play reinforcement learning update13. That such search is feasible at all, and moreover simple, is noteworthy, as test-time search in settings with as much hidden information as Stratego has been viewed as such a major technical challenge that previous work effectively forwent it1.
Ataraxos uses the belief network before each move to sample a collection of possible game states given its current position. For each candidate move, it runs depth-limited rollouts from these game states—starting with that candidate, then using its move network to simulate both players. Ataraxos estimates the value of the candidate move by averaging its network value predictions across the positions reached by rollouts starting from that move. These averaged values approximate self-play action values regardless of the policy of the opponent Ataraxos is facing, as the belief network approximates the posterior distribution of the self-play policy, the rollouts are executed by the self-play policy and the network predicts self-play position values.
Ataraxos uses these values to update its policy with a tabular step of magnetic mirror descent12, which regularizes and controls the size of the update with the same two reverse Kullback–Leibler divergences used during training. Importantly, this update can safely be more aggressive than those during training, both because the test-time update is tabular (and thus does not interfere with the policy at other positions), and because it is based on the more accurate advantage estimates enabled by the larger amount of computation per position at test time. The move Ataraxos plays is sampled from the updated policy. Set-ups, by contrast, are generated by sampling directly from the set-up network, with neither search nor any other intervention applied.
Training information
The reinforcement learning training run utilized 16 NVIDIA H100 graphics processing units (GPUs) for 1 week; the training run for the belief network utilized 4 such units for 4 days. Extended Data Fig. 2 shows metrics for this training run and searches performed on top of its final networks.
The reinforcement learning training run comprised 163 million finished games, 208 billion environment steps, 8.56 million gradient steps for the move network and 99.5 thousand gradient steps for the set-up network.
Compared with previous work1—which did not reach the level of top humans—our reinforcement learning training run consumed roughly 1/500th of the compute cost, 1/30th of the self-play games and 1/100th of the training examples. Taken together with the step change in playing strength reported in the main text, the reductions in self-play games and training examples indicate that the lower training cost reflects not only faster implementation but also markedly greater sample efficiency.
Rules of StrategoGame summary
Stratego is a board wargame played on a 10-by-10 grid with 92 occupiable squares and 2 blocks of non-occupiable squares called lakes (Extended Data Fig. 3a). Each player begins with 40 pieces, arranged in secret on the first 4 rows of their side so that the identities are concealed from the other player. The game proceeds in alternating turns. On each turn, the acting player moves one piece. If the piece is moved onto a square occupied by a piece of the other player, a battle occurs, at which point both pieces are revealed and at least one of them is removed from the board. Victory is achieved by capturing the Flag of the other player or by leaving the other player with no legal moves; the game ends in a draw if the player to move has no legal moves and the other player would have none were it their turn.
Piece details
A description of the pieces is provided in Extended Data Fig. 3b. Movable pieces—aside from Scouts—can move one square in a cardinal direction (that is, up, down, left or right) to squares that are either empty or occupied by an opponent piece; Scouts can move any number of squares in a cardinal direction to squares that are either empty or occupied by an opponent piece, so long as the movement does not jump over a lake or an occupied square. When pieces engage in combat, the outcome is typically determined by rank, which means that the higher-ranking piece defeats the lower-ranking piece if their ranks differ and that both are defeated if their ranks are the same.
Additional rules for competitive play
In competitive play, there are two additional rules. One is the two-square rule, which prohibits a piece from crossing the same square boundary on more than three consecutive turns of its owner. The other is the continuous-chasing rule, which sets limits on the ability of the pieces of one player to chase—in the sense defined below—those of the other.
Threats, evades, chases and chasing
We define a threat as an action that moves a piece adjacent to a piece of the opponent. An evade is an action that moves a piece that was threatened on the previous turn away from the piece that threatened it. A chase is an unbroken sequence of alternating threats and evades. Finally, chasing is the act of making threats during a chase.
The continuous-chasing rule states that a player who is chasing may not make a threat that would result in a position that has already taken place during the chase, unless that threat would return the moved piece to the square it occupied before the previous turn of the chasing player.
Additional rule for online play
There is also an additional rule implemented by Strategus, which is both the primary website for online competition and the website on which the evaluation against Pim Niemeijer was conducted. This rule is called the 200-move rule and states that the game ends in a draw if there is a sequence of 200 moves without a battle.
Time controls
Competitive Stratego also includes time controls. The evaluation against Pim Niemeijer was conducted under the default 15+3 Strategus time controls. The notation 15+3 means that each player starts with a 15-minute buffer and is allocated 3 free seconds for each move before their buffer starts to run down.
Engineering details
To make the project feasible on modest academic compute infrastructure, we implemented a graphics processing unit (GPU)-accelerated Stratego simulator in CUDA C++. This reduced the runtime of our data collection, training and search pipelines, while simultaneously reducing memory usage.
Our simulator was designed around five desiderata. First, it avoids explicitly storing in memory quantities that are easily recomputed. For example, we eliminate the need for a traditional rollout buffer that stores large information states and legal action masks by implementing a simulator capable of travelling back in the history of a game and reconstructing quantities of interest on-demand. This greatly reduces memory footprint and fragmentation, as discussed in the ‘Simulator and rollout buffer design’. Second, it maximizes simulation throughput, aiming for roughly 10 million board state updates per second, while also implementing anti-chasing rules. Third, it minimizes dynamic memory allocations: the simulator allocates memory as needed at construction time and then manages that memory directly rather than dynamically allocating and deallocating over time. Fourth, it supports search by enabling fast reset of board states to non-terminal states. Finally, it treats boards independently: once one game terminates, it is reset independently of whether the other games simulated in parallel have terminated. As a consequence, the boards gradually desynchronize, creating a distribution of training data that covers the different phases of the game.
Simulator and rollout buffer design
Unlike typical reinforcement learning infrastructure, we do not distinguish between the rollout buffer and the simulator. Rather, we expose a single object, called StrategoRolloutBuffer, that serves both purposes. This object tracks a tunable number N of parallel games and is responsible for two critical tasks: supporting historical queries (for example, returning a player’s legal action mask or information-state encoding at a past state); and ‘stepping’ the games by receiving actions and updating their states. The latter is performed by calling the ApplyActions method, which expects a tensor of N actions as input.
Once a game terminates, a new game starts. A new game is created by sampling a new initial board for both players. The distribution from which the initial board is sampled can be customized. It is also possible to ask the simulator to reset terminated games to a specific non-initial game state; this feature is important to support search, as discussed in ‘Reset behaviour and support for search’.
The rollout buffer tracks game states in a circular, preallocated GPU-memory buffer of tunable length. Correspondingly, queries about past states can be supported only if the past state is recent enough.
We found that this design, which integrates aspects of a traditional rollout buffer with a simulator, substantially reduces both memory fragmentation and usage compared with implementing a separate rollout buffer (for example, using PyTorch). A separate rollout buffer may need to make copies of legal action masks and information-state tensors, and pack them into aggregate tensors allocated by that buffer. Instead, our StrategoRolloutBuffer is able to reconstruct on-demand past information states and legal action masks. The training loop simply needs to remember at what historical time step the quantity needs to be computed, and then ask the backend to materialize the appropriate tensor to query the value and policy networks. Furthermore, this design is natural when considering that a Stratego simulator needs to track history (at least to a non-trivial extent) to implement the anti-chasing rules described in ‘Rules of Stratego’. Finally, by endowing the simulator with a notion of history, it becomes possible to implement efficient capturing of past states, as needed to efficiently implement search.
Two-square rule
To efficiently implement the rule, we built a custom state machine whose state is tracked and updated by StrategoRolloutBuffer.
The state machine is implemented internally by tracking the last four positions occupied by the last-acting piece for each player. The update logic runs directly on the GPU. Special care needs to be taken to properly account for Scouts, which have special movement abilities.
Continuous-chasing rule
To properly implement the rule, the simulator requires access to the board history of each game. To quickly detect whether a move would violate the rule, we used a state machine design paired with a fast diffing algorithm. Custom logic was added to make the rule compatible with resetting terminated games to start from non-initial states (as needed to support search; see ‘Reset behaviour and support for search’). Indeed, in this case, the continuous-chasing rule needs to be tested against the history of board states that leads to the non-initial reset state, rather than the history of boards stored in the rollout buffer.
Reset behaviour and support for search
To support search, the simulator needs to ‘pin’ the simulation of boards to start from a given state. This ability is implemented in our code by asking the StrategoRolloutBuffer to reset terminated boards not from an initial state, but rather from a custom non-initial state.
Further implementation details—including the simulator application programming interface (API), information-state representation, training configuration, network architectures and search hyperparameters—are provided in Supplementary Information.
Other experiments
We show reinforcement learning ablations in Extended Data Fig. 5 and the performance of search across varying hyperparameters in Extended Data Table 1. All ablations were performed directly from our final hyperparameter configuration, without retuning the remaining hyperparameters. The results therefore reflect the effect of each ablated component within our final design, rather than the best performance attainable under the ablated design choices.
We additionally compared the performance of the current iterates and their exponential moving averages over training across three seeds by evaluating each against a fixed reference model. As shown in Extended Data Fig. 6, the exponential moving average achieved similar or better mean performance while exhibiting lower variation across seeds.
We also evaluated Ataraxos against the bots from the Stratego Evaluator benchmark45: Asmodeus, Celsius, Celsius1.1, Vixen and Peternlewis. Ataraxos won 99, 98, 97, 99 and 98 of 100 games against these opponents, respectively. Win percentages this close to 100 imply that the Stratego Evaluator benchmark is meaningful only as a sanity check.
Learned set-ups
The probability distribution for each piece in a set-up chosen by Ataraxos is given in Extended Data Fig. 7, assuming that the Flag is on the left side of the set-up. The probabilities for the right side are symmetric, as Ataraxos applies a left–right orientation randomization to its set-ups after sampling from the set-up network to enforce symmetry. Ataraxos uses a bombed-in Flag (that is, a Flag that is enclosed by Bombs, such as in Extended Data Fig. 3a) about two-thirds of the time. The set-ups Ataraxos used in the series against Pim Niemeijer were generated fully autonomously by sampling from this distribution, one per game, and are shown in Extended Data Fig. 8.
Play style
The play style of Ataraxos differs from those of top human players (which are discussed, for example, by de Boer46), in terms of both set-ups and gameplay.
Set-ups
Players observed that Ataraxos uses aggressive set-ups (that is, ones in which high-value pieces are at or near the front), high Bombs (that is, Bombs in the third and fourth rows), and back-corner Flags (which humans consider difficult to defend) more frequently than humans. They also observed that it generally uses set-ups that have less predictable structure than those of humans.
Gameplay
Players observed the following. Ataraxos moves in a manner that makes its pieces hard to read (that is, discern the types of) relative to the movement patterns of human players. Ataraxos is more willing than humans to accept a draw during the opening when progressing the game would have negative expected value (as happened in Game 11 against Pim Niemeijer). Ataraxos has a stronger preference than humans for preserving Scouts deep into the game. Ataraxos uses certain bluffs sparingly relative to humans, but uses other bluffs that humans consider altogether too risky to play. Ataraxos fights more bitterly than humans when it is behind, aggressively stalling and pestering to slow the progression of its opponent—tactics some humans consider rude. Ataraxos excels relative to humans at long-term positional play, punishing mistakes, defending, playing from an information deficit, transitioning from middle- to end-game, and negotiating favourable positions into wins and unfavourable positions into wins or draws. Ataraxos takes gambles that humans consider arrogant (in the sense of not respecting the opponent). Ataraxos feels preternaturally lucky, always seeming to have the pieces it needs in the right places, to have its gambles pay off and to have its opponents do as it wants.
Contextualization vis-à-vis DeepNash
DeepNash1 is an AI for Stratego that was developed by DeepMind.
Evaluation
DeepNash was evaluated on the website Gravon in April 2022, winning 42 of the 50 games counted by Perolat et al.1, but not achieving the top ranking on the site. Several aspects of this evaluation are relevant to interpreting its results. (1) By the time of the evaluation, the player base had largely moved away from Gravon (only 25 players are listed for the final ranking of 202247)—as a result, the opponents matched against DeepNash were far from the level of top humans48. (2) The one-off online match setting may not have elicited the maximum effort level from the human opponents, who were not aware that an official evaluation was taking place. (3) These opponents had no reason to look for anti-bot exploits, as they were not aware that they were playing against a bot. (4) These opponents were not aware that DeepNash played a fixed strategy (and thus did not know that it was safe to exploit the same weakness across multiple games).
Separately, at the 2023 Stratego World Championship, DeepNash was demoed against human players, recording 19 wins and 9 losses49; DeepNash lost to most of the highest-ranked players who played against it, including Pim.
We reached out to DeepMind to ask whether they would allow an evaluation between DeepNash and Ataraxos, offering to build any infrastructure necessary for the evaluation. DeepMind responded that it would not be possible as the code for DeepNash is no longer functional.
Compute cost
Perolat et al.1 state that DeepNash was trained on 1,024 tensor processing unit nodes. To the recollection of the corresponding author of DeepNash with whom we spoke, this training took between 2 and 3 months and used tensor processing unit v3s. Under 2025 pricing50, such a training run would cost roughly between US$3,000,000 and US$4,500,000, depending on how much of the third month was used.
The reinforcement learning models and belief models of Ataraxos were trained on 16 H100s for 1 week and 4 H100s for 4 days, respectively. Such a run costs less than US$8,000 at 2025 prices51.
Sample cost
The training run for DeepNash consumed about 5.5 billion games and between 5 trillion and 10 trillion training examples.
The training run for Ataraxos consumed about 160 million games and about 50 billion training examples.
Opportunities for further improvementLearning
Ataraxos accesses history through features rather than learning across time directly. For the belief model, we found that much stronger compute-normalized performance could be attained by interleaving spatial attention with either temporal attention52 or recurrent models. However, for reinforcement learning, we did not observe an analogous out-of-the-box compute-normalized performance improvement owing to the additional runtime and memory requirements of such architectures, as well as their interplay with advantage filtering. We believe such architectures could achieve stronger performance given a sufficiently large compute budget or provided with additional runtime, memory or design optimizations.
Search
The amount of improvement attainable by the search procedure of Ataraxos is ultimately bounded because it is mimicking a single update step. A more sophisticated search algorithm would be able to leverage arbitrary amounts of additional compute to continue to improve the policy. One possible route towards this end would be to incorporate innovations from knowledge-limited subgame solving53.
Barrage Stratego evaluation details
The evaluations took place on Strategus using the default 5+1 time controls (meaning that each player starts with a 5-minute buffer and is allocated 1 free second for each move before their buffer starts to run down). Games were scheduled by the players at their convenience. Before the start of the evaluation, the players were informed that Ataraxos would not adapt to their play and that they would be evaluated by effective win rate.
Because belief-model training was ongoing at the time at which the evaluation began, we first evaluated the policy network (without search). The policy network won three series: 29 wins, 17 losses and 4 draws (a large margin) against world number-4 Axel Hangg; 36 wins, 12 losses and 2 draws (a very large margin) against world number-3 Sébastien Crot; and 26 wins, 21 losses and 3 draws against world number-1 Pim Niemeijer. We evaluated the search policy in a second 50-game series against Pim, chosen because he had lost to the policy network by the smallest margin. This second series put the search policy at a disadvantage: by the time Pim played it, he had already accumulated information about the set-up policy, which the two policies shared, and about their otherwise similar playstyles. Nonetheless, the search policy substantially outperformed the policy network, winning the series with 31 wins, 14 losses and 5 draws.
Implementation and training details are reported in ‘Barrage Stratego implementation details’ in Supplementary Information.
Hanabi evaluation details
We evaluated Ataraxos on the two- to five-player variants of Hanabi, with {0, 1, …, N} players running search in the N-player variant to demonstrate its power as a scalable multi-agent search method. Training, network and search details are reported in ‘Hanabi implementation details’ in Supplementary Information.
The evaluation results are summarized in Extended Data Fig. 9, with detailed numbers listed in Supplementary Table 32. All values are mean ± standard error computed over 10,000 games for each condition. With search employed by all players, Ataraxos achieves the following average scores and percentages of perfect (25-point) games:
2 players: score 24.654 ± 0.007, perfect games 77.53% ± 0.42%
3 players: score 24.863 ± 0.005, perfect games 89.90% ± 0.30%
4 players: score 24.852 ± 0.005, perfect games 88.40% ± 0.32%
5 players: score 24.410 ± 0.009, perfect games 58.05% ± 0.49%.
In Extended Data Fig. 9a, we plot the performance in terms of average scores over 10,000 games and compare these results with previous state-of-the-art results from ref. 54 for the two-player variant and ref. 55 for the three-, four-, and five-player variants. Importantly, performance improves monotonically as the number of players using search increases, with particularly large gains for variants with more players as a result of applying Ataraxos search to multiple agents. The gains in the average scores may seem numerically small; however, because Hanabi scores are capped at 25 points and additional points become increasingly difficult to secure near this ceiling, these results reflect substantial progress. Another way to understand this progress is through the percentage of perfect games that reach 25 points. In Extended Data Fig. 9b, we observe large increases in the percentage of perfect games across variants as we apply Ataraxos search to more players.
Dou Dizhu evaluation details
We further evaluated Ataraxos on dou dizhu. The reinforcement learning, belief-model and search details are reported in ‘Dou Dizhu implementation details’ in Supplementary Information.
We compared Ataraxos against the previous state-of-the-art method, PerfectDou23, as well as its predecessor, DouZero24. We evaluated heads-up performance in a duplicated setting: one agent plays as the landlord and the opposing agent controls the two independent peasants. Performance is reported as role-averaged score, computed by comparing the same pair of agents after swapping their roles. We evaluated Ataraxos against each baseline over 10,000 duplicated deals and report the performance in Supplementary Table 33. All values are mean ± standard error; comparisons involving Ataraxos use 10,000 duplicated deals, whereas comparisons among the policy network, PerfectDou and DouZero use 100,000 duplicated deals. The role-averaged scores of each agent against the others are as follows:
Ataraxos: 0.107 ± 0.012 versus the policy network, 0.199 ± 0.015 versus PerfectDou, 0.350 ± 0.016 versus DouZero
Policy network (i.e., Ataraxos without search): −0.107 ± 0.012 versus Ataraxos, 0.152 ± 0.004 versus PerfectDou, 0.286 ± 0.004 versus DouZero
PerfectDou: −0.199 ± 0.015 versus Ataraxos, −0.152 ± 0.004 versus the policy network, 0.142 ± 0.004 versus DouZero
DouZero: −0.350 ± 0.016 versus Ataraxos, −0.286 ± 0.004 versus the policy network, −0.142 ± 0.004 versus PerfectDou.
Ataraxos outperformed existing methods by a large margin and achieved the strongest performance against each baseline, establishing a new state of the art.
As a supplementary analysis, we evaluated Ataraxos under different update step sizes using the role-averaged score against PerfectDou. We report the results in Supplementary Fig. 8. To reduce evaluation cost, this sweep was run with a lighter schedule of 2,500 duplicated deals per setting.
Ethics statement
The Carnegie Mellon University Institutional Review Board (CMU IRB) determined that the human-player evaluation study qualified for exemption under the 2018 Common Rule, 45 CFR 46.104(d)(3)(i)(B), as a low-risk benign behavioural intervention (STUDY2025_00000013; modification MOD202500000509). Informed consent was obtained from all participants in accordance with the CMU IRB-reviewed consent procedures.