Skip to content

The Long Game Project · Reports · TLDR

We run wargames with AI players on live strategic questions, then publish what happened with the working shown.

What happens hereA strip of 3 numbered steps, in order: 1, Commission, A real question; 2, Play, AI players under a referee; 3, Present, A report you can argue with.1CommissionA real question2PlayAI players under a referee3PresentA report you can argue with
What happens hereA strip of 3 numbered steps, in order: 1, Commission, A real question; 2, Play, AI players under a referee; 3, Present, A report you can argue with.1CommissionA real question2PlayAI players under a referee3PresentA report you can argue with
Commission, Play, Present. Every report takes the same three steps, and a machine plays only the middle one.

The reports

Scenario exercises, 2026 to 2030, told as a timeline. Each one says how much a machine played. Read one in ninety seconds or in twenty-five minutes.

A free tool

A 54-question diagnostic that reads how your organisation decides. No AI scores it, and no account is needed.

Rehearsals, not forecasts

A run is a structured argument. If it sharpens your questions, it worked. If it tells you what to do, we wrote it badly.

Switch to Normal for the full page, or Deep for the working.

Wargames with AI players, on questions worth arguing about.

We take a live strategic question and build a matrix game around it. In a matrix game each player argues for a move and a referee rules on whether it works. AI players take the seats. Then we publish the run: the evidence pack of sources it drew on, the referee’s calls, and the source behind every number.

Rehearsals, with the working shown. Not forecasts.

Game design by Dr Dan Epstein, a medical doctor with a PhD, and director of the Long Game Project, where he has designed more than 130 business and strategy wargames. The machine plays. He is the author and reviewer.

Latest runs

Prussian-blue graded 19th-century anatomical plate of larynx figures, sun-bleached in a band reaching down from the top

The Throat to Choke

What AI does to the people who sell thinking, 2026-2030

The question: AI now does the junior work that accounting and consulting firms bill by the hour. Do the big firms protect that pricing model, or undercut it themselves first?

Read the report →

What this is. Wargames on live strategic questions. AI players, or programs following rules we wrote, take the seats, and a person or an AI referee rules on the moves. Each report says which on its front page. Where the engine played, the run has a seed, a fixed starting number, so the dice fall the same way when we re-run it. Every sourced number cites the evidence pack or a public source, and the build fails when a citation points at nothing. A number the author wrote is marked as his. Where the author, Dr Dan Epstein, overruled the machine, the report says so.

What this is not. Evidence about the future. A run is a structured argument. If a report sharpens the questions you are asking, it has done its job. If it tells you what to do, we have written it badly.

In one glance

How a report gets madeFlow diagram. 5 boxes: A question; A game with rules; AI players under a referee; A run on the record; A report, three depths. A question leads to A game with rules. A game with rules leads to AI players under a referee. AI players under a referee leads to A run on the record. A run on the record leads to A report, three depths.A questionA game with rulesAI playersunder a refereeA run onthe recordA report,three depths
How a report gets madeFlow diagram. 5 boxes: A question; A game with rules; AI players under a referee; A run on the record; A report, three depths. A question leads to A game with rules. A game with rules leads to AI players under a referee. AI players under a referee leads to A run on the record. A run on the record leads to A report, three depths.A questionA game with rulesAI playersunder a refereeA run onthe recordA report,three depths
Commission, Play, Present. The Long Game Project is Dr Dan Epstein's strategic-exercise practice, and this studio is part of it.

A free tool from the studio

The diagnostic

Fifty-four questions on how your organisation decides.

It takes about eighteen minutes, and no AI scores it. You get fifteen Ingredient scores, the things we measure, and a match to one of sixteen Flavours, the types we sort organisations into. You also get the blind spots, where your answers about what you value and your answers about what you do come apart.

Take the diagnostic

Commission a report on your question

Each report here is the version we can build without a client in the room. Public sources, our framing of the question, a machine at the table. A commissioned report uses your actors and your evidence, and nothing gets published. A tabletop puts your own team in the seats with the engine as referee.

Send the question in a sentence or two. We reply with a scope and a price in writing. Commissioned reports start at A$15,000 and tabletops at A$12,000. You decide.

The annex · Deep

What the front page leaves out: how much of the game behind each report a machine played, and how many times each one ran. A checked claim is a real-world fact in the report’s evidence pack, with a publisher, a quote, a retrieval date and a verdict, and a publication date where the source gives one. The check is made at retrieval, against the source.

No.TitleStatusFormatPlayedChecked claims
001The Throat to ChokePublishedOne story, written from 2030 looking backAn AI played the five roles one at a time. A person judged every move58
002The Seat PricePublishedOne game played by AI, then the same game re-run one hundred timesFour AI players, an AI referee, no person in the room. Then one hundred re-runs where only the dice changed16
005The Silence in the DoctrinePublishedForty-five games, fifteen per Treasury plan, played by rulesSeven roles, each a program following rules we wrote. Forty-five games, plus a check run with one assumption switched off55
006Written by the AdversaryPublishedForty-five games, fifteen per standard of proof, played by rulesSeven roles, each a program following a script we wrote. Forty-five games, plus a check run with one assumption switched off54
007Measuring the Strategic Stance of AI ModelsPublishedThree measurement studies: a repeated self-report by 17 models, then a peer-perception study and a behavioural panel on 8 flagships, 30 scenarios, 5 blind judgesThe models described themselves and each other, decided thirty cases, and a blind panel of models scored the decisionsNo pack yet
008The Quiet Spec ChangePublishedThree hundred distinct games, fifteen per route under each reading of the plan, played by rulesSix roles, each a program following rules we wrote. Sixty main games, replayed under four other sets of rules, 300 in all86

How many times each report ran

Report 001 ran once, with a person refereeing every move and no dice. Report 002 ran once with AI players and an AI referee. A fixed-rules version then ran one hundred times, with scripted players and no AI referee, and only the random draws changed. Reports 005 and 006 each ran forty-five games: fifteen per Treasury plan in 005, fifteen per standard of proof in 006, plus a check run with one assumption switched off. That is enough to see which outcomes are fragile to the dice and which are not. Monte Carlo, the method that estimates the odds of an outcome from many random trials, needs far more runs than that, and we do not claim it.

The five rules every report follows are stated once, on Methods, under “The rules we do not bend”. The evidence for whether the players are any good is on the evidence page.