Skip to content

Strategy Soup · The Long Game Project · TLDR

We run wargames with AI players on live strategic questions, then publish what happened with the working shown.

What happens hereA strip of 3 numbered steps, in order: 1, Commission, A real question; 2, Play, AI actors under a referee; 3, Present, A report you can argue with.1CommissionA real question2PlayAI actors under a referee3PresentA report you can argue with
What happens hereA strip of 3 numbered steps, in order: 1, Commission, A real question; 2, Play, AI actors under a referee; 3, Present, A report you can argue with.1CommissionA real question2PlayAI actors under a referee3PresentA report you can argue with
Commission, Play, Present. Every report takes the same three steps, and a machine plays only the middle one.

The reports

Scenario exercises, 2026 to 2030, told as a timeline. Each one says how much a machine played. Read one in ninety seconds or in twenty-five minutes.

Two free tools

A 54-question diagnostic that reads how your organisation decides, and a twin, an AI that answers as your organisation would. No account needed.

Rehearsals, not forecasts

A run is a structured argument. If it sharpens your questions, it worked. If it tells you what to do, we wrote it badly.

Switch to Normal for the full page, or Deep for the working.

Strategy Soup · Exercised scenario reports · The Long Game Project

Synthetic wargames on questions worth arguing about.

We take a live strategic question and build a matrix game around it. In a matrix game each player argues for a move and a referee rules on whether it works. AI-played actors take the seats. Then we publish the run: the evidence pack of sources it drew on, the referee’s calls, and the source behind every number.

Rehearsals, with the working shown. Not forecasts.

What this is. Wargames on live strategic questions. AI actors play them, and a human or an AI arbiter referees. Each report says which on its front page. Where the engine played, the run has a seed, a fixed starting number, so anyone can re-run it. Every number cites the evidence pack or a public source, and the build fails when one does not. Where the author, Dr Dan Epstein, overruled the machine, the report says so.

What this is not. Evidence about the future. A run is a structured argument. If a report sharpens the questions you are asking, it has done its job. If it tells you what to do, we have written it badly.

In one glance

How a report gets madeFlow diagram. 5 boxes: A question; A game with rules; AI actors under a referee; A run on the record; A report, three depths. A question leads to A game with rules. A game with rules leads to AI actors under a referee. AI actors under a referee leads to A run on the record. A run on the record leads to A report, three depths.A questionA game with rulesAI actors undera refereeA run onthe recordA report,three depths
How a report gets madeFlow diagram. 5 boxes: A question; A game with rules; AI actors under a referee; A run on the record; A report, three depths. A question leads to A game with rules. A game with rules leads to AI actors under a referee. AI actors under a referee leads to A run on the record. A run on the record leads to A report, three depths.A questionA game with rulesAI actors undera refereeA run onthe recordA report,three depths
Commission, Play, Present. The machine plays. Dr Dan Epstein writes the argument. Long Game Project is his strategic-exercise practice, and this studio is part of it.

Latest runs

Prussian-blue graded 19th-century anatomical plate of larynx figures, sun-bleached in a band reaching down from the top

The Throat to Choke

What AI does to the people who sell thinking, 2026-2030

The question: AI now does the junior work that accounting and consulting firms bill by the hour. Do the big firms protect that pricing model, or undercut it themselves first? Neither: the client takes the fee back, and the fight that decides the outcome is over who carries the blame when the work is wrong.

One story, written from 2030 looking back · September 2026 · An AI played the five roles one at a time. A person judged every move · 58 checked claims: real-world facts, each with a source and a check date

Read the report →

002 · Published

The Seat Price

The question: Business software is mostly sold one licence per person, and AI agents now do the work those people did. So when does the price get reset, and who pays for it? In the games the published list price held. What broke was the secret of what each customer really pays.

One game played by AI, then the same game re-run one hundred times · September 2026 · Four AI players, an AI referee, no person in the room. Then one hundred re-runs where only the dice changed

005 · Published

The Silence in the Doctrine

The question: Washington agrees the Treasury, not the Fed, decides how much of the national debt is long-term. Nobody has said whether the Treasury may use that power to push long-term interest rates down. Who forces that question into the open, and what does the reason the Treasury gives on 4 November 2026 cost it?

Forty-five games, fifteen per Treasury plan, played by rules · September 2026 · Seven roles, each a program following rules we wrote. Forty-five games, plus a check run with one assumption switched off

006 · Published

Written by the Adversary

The question: In July 2026 the only account of what a group of AI agents was doing was the one the agents wrote themselves. What should such an account be allowed to prove, and who decides?

Forty-five games, fifteen per standard of proof, played by rules · September 2026 · Seven roles, each a program following a script we wrote. Forty-five games, plus a check run with one assumption switched off

003 · On the slate

The First Hundred Days

The question: After the deal closes, who kills whose system, and how long does the losing side keep running it anyway?

Played by people at a table, with the engine as referee · Not yet played

004 · On the slate

The Resilience Game

The question: When a shock lands on a supply chain nobody owns end to end, who moves first, and who waits for someone else to?

Played by people at a table, with the engine as referee · Not yet played

Two free tools from the studio

The diagnostic

Fifty-four questions on how your organisation decides.

It takes about eighteen minutes, and no AI scores it. You get fifteen Ingredient scores, the things we measure, and a match to one of sixteen Flavours, the types we sort organisations into. You also get the blind spots, where what you say and what you do come apart.

Take the diagnostic

The twin

Your persona, talking back.

Load the persona the diagnostic wrote and ask it a strategic question. It answers in character and tells you when it had to stretch. Free while it lasts. You bring your own API key, the key your AI provider gives you.

Talk to your twin

Every report goes through the same three rooms: Commission, Play, Present. A human picks the question and signs a brief with kill criteria, the conditions that stop a run before we publish it. Some runs stop there. The engine plays it: actors claim, rivals object, an arbiter rules with reasons, and a die roll settles whether the move works. The engine rolls that die, and no AI can touch it. Then Dr Dan Epstein reads the whole record and writes the argument. The machine plays. He is the author.

The full method

Commission a run with your team at the table

Each report here is the version we can build without a client in the room. Public sources, our framing of the question, a machine at the table. A commissioned run uses your actors and your evidence, your team plays the seats, and nothing gets published.

If one of these questions is yours, or should be, tell us. The button opens the Long Game Project contact page. Send the question in a sentence or two. We reply with a scope and a price. You decide.

Commission a report

The annex · Deep

What the front page leaves out: how much of each report a machine played, and how many times each one ran. A checked claim is a real-world fact in the report’s evidence pack, with a publisher, a date, a quote, a retrieval date and a verdict.

No.TitleStatusFormatPlayedChecked claims
001The Throat to ChokePublishedOne story, written from 2030 looking backAn AI played the five roles one at a time. A person judged every move58
002The Seat PricePublishedOne game played by AI, then the same game re-run one hundred timesFour AI players, an AI referee, no person in the room. Then one hundred re-runs where only the dice changed16
005The Silence in the DoctrinePublishedForty-five games, fifteen per Treasury plan, played by rulesSeven roles, each a program following rules we wrote. Forty-five games, plus a check run with one assumption switched off55
006Written by the AdversaryPublishedForty-five games, fifteen per standard of proof, played by rulesSeven roles, each a program following a script we wrote. Forty-five games, plus a check run with one assumption switched off54
003The First Hundred DaysOn the slatePlayed by people at a table, with the engine as refereeNot yet playedNo pack yet
004The Resilience GameOn the slatePlayed by people at a table, with the engine as refereeNot yet playedNo pack yet

How many times each report ran

001 ran once, with a person refereeing every move. 002 ran once with AI players and an AI referee, then one hundred more times under fixed rules where only the dice changed. 005 and 006 each ran forty-five games: fifteen per Treasury plan in 005, fifteen per standard of proof in 006, plus a check run with one assumption switched off. That is enough to see which outcomes survive a change of dice and which do not. It is not thousands of runs, so we do not call it Monte Carlo, the method that runs thousands of random trials.

The five rules every report follows are stated once, on Methods, under “The rules we do not bend”. The evidence for whether the players are any good is on the evidence page.