Ataraxos Beats Top Stratego Players 15-1-4, Using Under 1/100 of DeepNash Training Data
Updated
Updated · MIT News · Sep 30
Ataraxos Beats Top Stratego Players 15-1-4, Using Under 1/100 of DeepNash Training Data
2 articles · Updated · MIT News · Sep 30
Summary
MIT-led Ataraxos posted a record 15-1-4 against the world’s strongest Stratego player and went 39-2 versus top humans at the Stratego world championship, marking the first clear superhuman result in the hidden-information game.
Less than 1/100 of DeepNash’s training examples and less than 1/30 of its self-play games were enough because the system pairs self-play reinforcement learning with decision-time planning that estimates opponents’ hidden pieces before each move.
10^66-plus possible piece configurations make Stratego far harder to brute-force than chess, and earlier systems including DeepMind’s remained too costly and too weak to beat elite human players.
The researchers also adapted Ataraxos to Barrage Stratego, Hanabi and Dou dizhu, where it again reached superhuman performance, suggesting broader use in negotiations, cybersecurity and military planning.
Nature published the work as the team said the next step is interpretability, so humans can audit and overrule the model before any real-world deployment.