Scientists from the Massachusetts Institute of Technology (MIT), Carnegie Mellon University, New York University and Stanford have unveiled an artificial intelligence (AI) system that beat the world's strongest Stratego players by a decisive margin. According to MIT News, the system, called Ataraxos, finished its match series against the world champion with 15 wins, 1 loss and 4 draws, and beat the leading players at the world championship 39-2. The research was published in the journal Nature. No AI system had ever achieved such a result before. Stratego is a two-player board wargame in which the identity of the opponent's pieces stays hidden throughout the game.

Record Result

To put it to a practical test, the researchers took Ataraxos to the world championship. The system finished its series against the strongest player in Stratego history 15-1-4 β€” the largest margin ever recorded against an opponent of that caliber. The 20-game series produced 15 wins, 1 loss and 4 draws. The total against the championship's other leading players was 39-2 β€” the system won 39 of 41 games. According to MIT News, no AI system had previously beaten the strongest human Stratego players by such a decisive margin.

Scientists from four universities β€” MIT, Carnegie Mellon, New York University and Stanford β€” worked on the research together. The fact that the system beat not only the champion but also a group of the championship's strongest players shows the win was no fluke. Ataraxos is thus the first AI system in history to defeat a world champion by a decisive margin.

How the System Works

Ataraxos was trained with reinforcement learning β€” through self-play: the system played against itself many times, developing a base strategy that confers an advantage in Stratego. The researchers developed algorithms that made training significantly more efficient β€” as a result, the system reached championship level on a comparatively small training budget. As Ars Technica notes, the system beat the strongest player in history on a comparatively small training budget.

During play, Ataraxos uses the base strategy as a starting point but refines its choice on the spot before every move. This uses a method called decision-time planning: generative AI estimates through probabilities which pieces the opponent has hidden, then analyzes future moves and picks the best one. According to the researchers, this method is what let the system surpass human performance.

The system's training efficiency was directly compared with DeepMind's previous DeepNash system: Ataraxos reached strictly higher playing strength β€” requiring less than one hundredth of the training examples and less than one thirtieth of the self-play games. In other words, the system is not only stronger than the previous record holder but also took significantly fewer resources to train. Lead author Gabriele Farina put it this way:

"The system reached a strictly higher playing strength than DeepMind's DeepNash β€” requiring less than one hundredth of the training examples and less than one thirtieth of the self-play games," said Gabriele Farina.

The study, published in the journal Nature, details these efficiency results, the championship test conditions and the comparative analysis with DeepNash.

Winning at Other Games Too

To show that the method works not just for Stratego but in general, the research team adapted Ataraxos to three other imperfect-information games: fast-paced Barrage Stratego, the multiplayer cooperative card game Hanabi, and Dou dizhu, in which two players play against a third. The system achieved superhuman performance in each one. As MIT News writes, these results confirm the method's versatility across different settings. The results and the system's training details are detailed in the scientific study published in Nature.

Why This Achievement Matters

Stratego is an imperfect-information game: the identity of the opponent's pieces stays hidden until they collide. Players place pieces on their half of the board and try to capture the opponent's flag. Imperfect-information games have long been an unconquerable test for AI β€” Stratego is just such a game. That is why it is considered an important benchmark for testing AI systems' strategic reasoning. As Ars Technica emphasizes, this hidden-information game long stumped AI systems, while Ataraxos beat it on a comparatively small training budget. The result was also covered in Ars Technica's October 1 story.