Stratego board with hidden opponent pieces illustrating AI decision-making under uncertainty

DeepMind threw an enormous budget at Stratego and still couldn’t beat the best humans. A team from Carnegie Mellon, MIT, NYU, and Stanford just did it with 16 GPUs and a few thousand dollars. Their system, Ataraxos, beat Pim Niemeijer, arguably the strongest Stratego player ever, 15 games to one. The difference wasn’t compute. It was a second neural network, a belief model, trained to guess what the opponent’s hidden pieces actually were based on how they’d been played.

That is the same shape of problem you deal with every week. Incomplete sensor coverage, operators who log some things and not others, quality data that arrives three shifts late. Below, what the Ataraxos approach to AI decision-making under uncertainty actually means for how you build systems on plant data you don’t fully trust.

Chess Fell in 1997. Stratego Held Out Until Last Week

Deep Blue took Kasparov in 1997. AlphaGo took Lee Sedol in 2016. Poker bots have been beating professionals for years. Stratego stayed unsolved, and DeepMind’s DeepNash, released in 2022, did not close it.

The reason is structural. Forty pieces in any order produce more than a decillion possible setups. A single game can run 2,000 moves against roughly 40 in chess. And the hidden information leaks out slowly across that entire span, so you are making consequential moves for hours on partial knowledge that only resolves in fragments.

Read that description again without the board. Enormous state space, long horizons, and most of what matters unobserved until something collides. That is a plant floor. Ataraxos is the first credible demonstration of AI decision-making under uncertainty at that scale, which makes it an operations result wearing a board game costume.

Timeline of chess, Go, and poker milestones leading to AI decision-making under uncertainty in Stratego

What Ataraxos Actually Achieved, And What It Cost

The match record against the best human player of all time

Twenty games. Fifteen wins for the machine, one for the human, four draws. That is not a narrow edge or a statistical wobble that disappears on a rematch. Against a player of Niemeijer’s calibre, it is a decisive gap in playing strength.

The margin matters more than the headline. Earlier systems could play respectable Stratego and lose to top humans in ways that looked close. Ataraxos did not squeak through. It produced a record that leaves no argument about whether the approach works, which is exactly the standard you should hold any AI deployment to before it touches a production decision.

16 GPUs and a few thousand dollars versus DeepMind’s budget

Training ran across 163 million self-play games on 16 GPUs for a few thousand dollars. DeepMind, with what the researchers describe as an exceptional budget, could not get DeepNash past the best humans. The team that did spent less than most manufacturers spend on a single consulting engagement.

Two things made the difference, and neither was hardware. The first was pacing the learning: big, bold strategy changes early in training, small careful ones later, because hidden information sends self-play algorithms in circles. The second was the belief model that let the system search ahead before each move, something DeepMind had written off as intractable.

Factor DeepNash (2022) Ataraxos
Budget Exceptional A few thousand dollars, 16 GPUs
Lookahead search None Yes, via belief model
Result vs top humans Did not reliably win 15-1-4

Take the first practical lesson into your next vendor conversation. When someone sells you compute, model size, or parameter counts as the reason their system will handle AI decision-making under uncertainty on your plant floor, they are describing the wrong constraint. Problem formulation is the binding constraint. Ask how the system reasons about what it cannot observe.

The Belief Model: How Ataraxos Reasons About What It Cannot See

Inferring hidden pieces from observed behaviour

The belief model is a separate neural network with one job: estimate what each unknown enemy piece actually is, based on how that piece has been played. A piece that advances into contested space early behaves differently from a bomb sitting next to a flag. Those behavioural traces are data, and the model converts them into a probability distribution over identities.

That conversion is the whole trick. An unobservable board becomes a scored set of possibilities the system can reason over and act on. It does not need certainty about any single piece. It needs a calibrated estimate it can weigh against the cost of being wrong, which is exactly what AI decision-making under uncertainty requires in practice.

Bluffing is why this is hard. A weak piece pushed forward aggressively is meant to read as a marshal. Any model that takes observed behaviour at face value gets exploited. Ataraxos has to price in deception as part of the distribution, not treat it as noise.

Why search before each move needed a belief model first

AlphaGo-style systems refine strategy with a search immediately before acting. DeepNash never did this in Stratego because the search space was judged too large to make it work. That left an open question, and Gabriele Farina put the answer bluntly:

This is one of the things that we did figure out how to do.

The belief model is what made the search tractable. You cannot look ahead across a decillion possible setups. You can look ahead across the handful of configurations the belief model rates as plausible. Narrow the state space first, then search it.

Training schedule mattered too. Hidden information sends self-play algorithms around in circles, so the team made big, bold strategy changes early in training and small, careful ones later. Across 163 million self-play games, that annealing is what stopped the system from cycling instead of converging.

Diagram of a belief model network estimating hidden states for AI decision-making under uncertainty

Your Plant Data Is a Stratego Board, Not a Chessboard

Your plant reveals itself on collision. A root cause surfaces when a batch fails final inspection. A bearing announces itself when it seizes. A supplier’s process drift becomes visible when their material behaves badly in your line. Until contact, you are looking at positions without identities.

That is the mechanic Eugene Vinitsky described as “a massive amount of hidden information that unfolds over a very long time scale.” Sensors cover some stations and not others. Operators log some deviations and skip the rest. You never see your supplier’s internal process data at all.

Where perfect-information tooling quietly fails on the floor

Static dashboards assume the board is visible. They plot what was measured and leave the gaps blank, which trains people to treat unmeasured as fine. Fixed rule thresholds do the same thing in a more dangerous way: they fire on a single observed variable and stay silent about everything the sensor cannot reach.

The failure is quiet because nothing looks broken. Your SPC charts are in control. Your OEE report is green. Then a defect cluster appears three weeks later with no trail, because the information that would have explained it was never represented anywhere in your tooling.

What an explicit belief estimate changes about a decision

Ataraxos did not solve Stratego by seeing more. It solved it by holding a scored estimate of what it could not see and updating that estimate with every move. The equivalent on your floor is a model that carries a live probability distribution over unobserved state.

Three forms of that are already practical. A predicted defect cause distribution that ranks likely contributors before teardown, so the investigation starts in the right place. A supplier risk posterior that updates on incoming inspection results, delivery behaviour, and certificate patterns rather than waiting for an audit. A machine condition estimate that interpolates between inspections instead of assuming health persists until the next check.

Each one turns a blank into a number you can act on. That changes the decision, not just the reporting.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.

What Operations Leaders Should Take From This Before 2027 Budgets Close

The thread worth watching is not the win record. It is that belief modelling plus targeted search attacked the uncertainty directly instead of adding parameters, and it did so on a budget any mid-sized manufacturer could approve. Cheap techniques diffuse fast. Expect this pattern in commercial tooling within two budget cycles.

Start by auditing your top ten recurring decisions and marking which ones are actually made with full observability. Scheduling with complete machine status is one problem. Releasing a batch when three of eight stations are instrumented is a different problem entirely, and most plants treat them identically.

Three questions to ask your next AI vendor

Vendors will happily sell you a model trained on the data you have. Few will tell you what that model silently assumes about the data you do not have. Ask directly:

  • What does your model assume about unobserved variables: does it ignore them, impute an average, or estimate a distribution over them?
  • Does it output confidence alongside the recommendation: a prediction without an attached belief state is a guess wearing a uniform.
  • How does it update when hidden information finally surfaces: when a failure reveals a root cause, does that revision propagate or get discarded?

Then size the pilot on problem clarity, not GPU budget. Ataraxos needed 163 million self-play games, but the real engineering was in the training schedule: bold strategy changes early, small careful ones later. As MIT’s Gabriele Farina put it, “This is one of the things that we did figure out how to do.” Figuring out beats spending.

Where the ROI actually lands

Better decisions under uncertainty do not show up as a line item called AI. They show up as fewer customer escapes, because marginal batches get held on probability rather than gut feel. They show up as less rework, because interventions land before the defect propagates downstream.

The largest and least measured return is time. Quality teams spend enormous hours reconstructing what happened after the fact. A system that maintains explicit beliefs about unobserved conditions has already done most of that reconstruction, continuously, before anyone opens an investigation.

Source: arstechnica.com

Leave a Reply