Deep Blue took down Garry Kasparov at chess in 1997, AlphaGo beat Lee Sedol at Go in 2016, and poker bots have been beating professionals for years. But one classic game called held out. Even DeepMind, with its exceptional budget, couldn’t build a machine that reliably beat the best human players.
Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best player of all time, 15 games to one, with four draws. And it took just 16 GPUs and a few thousand dollars to train it.
Hidden armies
In , each player gets 40 pieces representing military ranks, from a marshal down to a spy, plus bombs and a flag. You win by capturing the opponent’s flag. Your opponent knows your pieces are, but not they are. Identities are revealed only when two pieces collide in battle—the weaker one is removed, and the identity of the winner is revealed. That makes an imperfect-information game, just like poker, which computers cracked years ago. “There’s something super distinctive about , which is that it is a massive amount of hidden information that unfolds over a very long time scale,” said Eugene Vinitsky, a researcher at NYU and co-author of the study.
In some forms of poker, the hidden information is tiny. In Texas Hold’em, “You only have two hidden cards,” said Gabriele Farina, an MIT computer scientist and another co-author. That leaves just 1,326 possible hands, few enough for a machine to weigh them all. “In , there’s 40 pieces on the board that could be in any order,” Farina said. That’s more than a decillion possible setups. Then there’s the game’s length.
“In chess, usually the game lasts 40 moves, but in , a game can easily last 2,000 moves,” Farina said. On top of that, is a game of bluffing. Sometimes you move a weak piece as if it were a marshal, just to scare the opponent off. When players bluff too often, their threats mean nothing; when they never bluff, they become predictable. That balancing act, the team explains, is what stumped earlier AIs like DeepMind’s DeepNash, introduced in 2022.
Learning to guess
Just like DeepNash, Ataraxos learned by playing against itself—163 million games in total. In these self-play sessions, moves that led to wins were reinforced and played more often in future matches, while moves that led to losses were played less, which was the same simple training idea. The difference was in how much Ataraxos adjusted after each game, because hidden information tends to send self-play learning algorithms around in circles. The team addressed this by making big, bold changes in strategy early in training and small, careful ones later.
The even bigger innovation was something DeepNash never had: thinking ahead before each move. AIs like AlphaGo refine their general strategy with a search just before acting. DeepMind couldn’t make that work in because the search space was too large, leaving it an open question whether it was worth trying.
“This is one of the things that we did figure out how to do,” Farina said. The solution was a second neural network, a belief model, trained to guess the opponent’s hidden pieces based on how they had been moving. This way, instead of iterating through every possible arrangement, Ataraxos samples plausible ones, plays out candidate moves in each, and picks based on how they turned out.
And it shows in its playstyle.
Calm and unbothered
The name Ataraxos comes from the ancient Greek word for a state of calm. “It means somebody that’s calm and unbothered,” Farina explained. He suggests the structure of the AI and its lack of human emotions ensure it doesn’t react impulsively, “even in situations where a human would be losing their mind.” While the human might try big gambles to come back from a significant deficit, Ataraxos would work its way back into the game slowly and methodically.
The strategy it developed also avoids drawing attention to any problems it faces. When Ataraxos estimates its opponent has no reason to suspect a weak spot, it leaves that spot alone, even if it might look like a disaster waiting to happen to anyone who can see both sides of the board.
“For humans, it’s very hard when you know a secret to make decisions ignoring the fact that you know that secret,” Farina said. “For machines, it’s easy.” This, the team says, leads machines to make moves a human would only do while bluffing—and follow up on them much better than humans. “We would watch the bot ‘bluff’ its way back from like a two percent victory probability, very, very casually,” Vinitsky added.
Niemeijer, the human player Ataraxos pulled these miraculous comebacks against, has four world championships and more than 600 weeks as the world’s top-ranked player.
The match
Over three weeks, Niemeijer played 20 online games against Ataraxos, earning $100 for each win. He knew the AI would not adapt to him, which gave him time to hunt for weaknesses. He managed to win just once.
That loss, researchers claim, wasn’t really a flaw. Playing well requires randomizing the arrangement of your pieces, so luck always plays a role. “Even a perfect strategy, sometimes it will just lose,” Farina said.
The human champion apparently got lucky, but it went both ways. “Sometimes we got lucky,” Vinitsky admitted.
At the 2025 Stratego World Championship, attendees who challenged Ataraxos fared even worse. The AI won 38 of 40 games. In the process, it also changed how people play. “I think this bot has kind of skewed the metagame a little bit,” Farina said.
Players were surprised, for example, by how often it tucked its flag into a corner behind just two bombs, a rarely played setup.
But Ataraxos’s best trick was arguably its price tag.
The price to pay
DeepNash was trained for two to three months on 1,024 of Google’s specialized chips, a run the Ataraxos team estimates would cost $3 million to $4.5 million at 2025 prices. Ataraxos, in contrast, needed 16 GPUs for a week, plus an additional four GPUs for four days to train the belief model.
Farina and lead author Samuel Sokota achieved this efficiency by writing a simulator that runs millions of moves per second on graphics cards. “At the scale that we are in academia, we don’t really have access to an entire field of GPUs,” Farina said.
The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger.
The Ataraxos architecture also worked in learning other games. The same approach beat three world champions at , a faster eight-piece variant of , mastered the cooperative card game , and beat the best bots at the Chinese card game dou dizhu. But the team has its sights set on scenarios far more complex than board or card games.
Beyond the board
Board games have fixed rules and clear winners, while real-world problems like negotiations, financial markets, or military conflicts usually don’t. But the Ataraxos team argues the gap is smaller than it looks, since tackling any real problem starts with building a simplified model of it.
“War gaming is a common thing that people do,” Vinitsky said. “You can use the techniques that were derived here to play it forward and see how a strong opponent might respond to what you do.”
Farina and his colleagues are now interested in making their AI more understandable, since Ataraxos, in its current state, can’t explain why it makes the moves it makes. “We work on machines that produce strong but also interpretable and explainable strategies. I think we’re not quite there yet,” Farina said.
Nature, 2026. DOI: 10.1038/s41586-026-11036-y

