Skip to content

The Long Game Project · How it works · The engine · TLDR

A question, played as a game

A question becomes a game. Each party gets a seat and a private brief, its goals and limits. An arbiter, the referee, sets the odds on each move. Seeded dice, rolls that repeat exactly from a starting number, decide. Dr Dan Epstein writes the argument.

One beat of playA seat makes a claim, the other seats object, the arbiter rules, a die rolls, and all of it is recorded. Then the loop runs again.ClaimObjectRuleRollRecord
One beat, one turn of play: a claim, the objections, a ruling, a roll, a record. The game is this loop, run again and again.

The table

The seats are archetypes, types drawn to be recognisable, rather than real companies. Each has a goal, win conditions, red lines and a read on everyone else.

The beat

A seat argues for a move. The others argue against. The arbiter sets a target. A die settles it. All of it goes on the log.

One hundred worlds

For “The Seat Price” we re-ran the game one hundred times under fixed rules, with only the dice changed, so the report describes a spread of endings instead of one story. Each report states its own count.

Not a forecast

The engine knows nothing about 2030. It makes assumptions arguable. That is the whole job.

Switch to Normal for the full page, or Deep for the working.

The Long Game Project · Method · The wargame engine

What the wargame engine is, and what it is not

A question becomes a game. Each party gets a seat and a private brief, its goals and limits. An arbiter, the referee, sets the odds on every move. Seeded dice, rolls that repeat exactly from a starting number, decide. Dr Dan Epstein reads the transcripts and writes the argument. This page explains the machinery. It makes no claim that the machinery predicts anything.

An exercised scenario, in one paragraph

We take a question a leadership team is arguing about and turn it into a game with rules. Each party who matters gets a seat and a private brief: a goal, win conditions, red lines, and how it reads everyone else at the table. Play advances in beats: a seat argues for a move, the other seats argue against it, an arbiter sets the odds, and a die settles it. The log keeps every claim, argument, ruling and die face. Then Dr Dan Epstein reads the whole record and writes the argument. He marks where he overrode the machine and where the evidence cuts against the game.

What it is for. Making assumptions arguable. A room usually agrees on the frame before it starts arguing, and the frame is where the money is. Running the game forces the frame into the open. The incumbent keeps taking revenue from its old price model because of something nobody wrote down, or the challenger’s pitch fails for a reason nobody had named. Our published reports name their load-bearing assumptions, the ones an ending cannot stand without, so you can disagree with them by name.

What it is not. It is not a forecasting tool. The engine knows nothing special about 2030. The odds on a move come from a language model’s judgement, converted to a number on a table, then settled by a seeded die. Our reports say the same on their own front pages: scenarios, not forecasts. “The Seat Price”, one of the two published reports, says in its method note, the section that explains how it was played, that we make no claim the games predict real-world outcomes.

Who sits at the table

The seats are archetypes, not real companies. “The Seat Price” had four. Atlas is a megavendor whose revenue base is per-user licences, and Forge is an AI-native challenger that prices per outcome. Meridian is an enterprise buyer, and Pillar is a global systems integrator. Each started with a declared strategic position and played eight half-year moves, from 2026 H2 to 2030 H1. They are composites drawn to be recognisable, and the report says which real behaviour each one copies.

Where the personas come from. The diagnostic asks 54 questions, scores an organisation on 15 strategic dimensions, and places that score near one of 16 archetypes. When we design a game, we ground each seat in one of those archetypes, so it plays a consistent character instead of a generic role. That grounding goes into the seat’s private brief. The step from questionnaire to brief is the one we found broken in August 2026 and fixed. The validation page tells that story. Read it before you trust any of this.

The arbiter plays no seat. It rules on claims. Its job is to set a difficulty for a proposed move from the arguments for and against, and to stay consistent with its earlier rulings in the same run. Every ruling puts four things on the run’s event log: the decision, the reasoning path, the assumptions the ruling rests on, and a self-reported confidence. The engine never presents a ruling as certain.

How much rope the arbiter gets is a setting. The arbiter runs in one of four modes. It can rule at once and log the reasoning, or settle routine calls itself and halt for confirmation on major ones. It can propose every ruling for a human to approve. Or a human can rule while the arbiter assists. Which mode an exercise ran in is a fact about that exercise, and we state it per report.

How one beat works

A beat is the smallest unit of play. Every beat runs the same five steps.

  1. Each seat proposes its move as a claim and states the odds it thinks the move deserves. It also records a reasoning journal: what it considered, what it rejected, and how confident it is.
  2. Every other seat argues against that claim, in writing. This step is what makes it a game and not a self-graded essay. If a seat cannot argue, the engine warns that the other seats’ claims will face fewer objections.
  3. The arbiter reads the claim, the arguments for and against, and the seat’s current position. Then it sets a target number. Stated odds convert to that target on a fixed table. A move argued at 90 per cent becomes an easy target, and a move argued at 5 per cent becomes a hard one.
  4. A twenty-sided die rolls against the target, with modifiers from the seat’s resources added on. A natural 1, the die itself showing 1, always fails. A natural 20 clears any target of 20 or under. How far the total clears or misses the target sets how decisively the move landed.
  5. The engine appends the claim, the objections, the ruling, the die face and the outcome to the run’s event log.
The beat loopA seat makes a claim, the other seats object, the arbiter rules, a die rolls, and all of it is recorded. Then the loop runs again.ClaimObjectRuleRollRecord
The same five steps, every beat, until the game ends. Nothing in a report happens outside this loop.

The engine never stores game state directly. It stores events and rebuilds the state by replaying them. That is why a what-if is cheap: copy the log up to a point, change one thing, replay from there. It is also why you can audit a run beat by beat instead of taking a summary on trust.

One hundred re-runs turn a story into a spread

One playthrough is an anecdote. The dice could have fallen the other way, and a report built on a single run is a report built on luck. So for “The Seat Price” we also ran a fixed-rules version of the same game one hundred times: same seats, same menu of moves, different dice. Then we read the spread instead of the story.

  • The hundred runs are not the language models playing. Seeded policy agents, players that follow a fixed policy instead of a language model, drive the fixed-rules version. The model seats that played the live run do not drive it. The hundred runs tell you about the shape of the game we wrote. They tell you nothing about how a model behaves. “The Seat Price” sets this out in its own method note.
  • A spread instead of a story. The fixed-rules game scores each seat’s strategic position. Atlas, the incumbent, starts on 100, and a seat that reaches zero is out of the market. Across the hundred worlds Atlas finished anywhere from 82 to 112, with a median of 95. Those bounds are the 5th and 95th percentiles, so about nine worlds in ten ended between them. Not one world produced a collapse, meaning no run drove Atlas out of the market inside the window. The report says that is partly a property of the board: on these rules, falling from 100 to zero inside eight half-years needs close to worst-case dice on almost every turn.
  • We compare forks seed for seed. To test “what if the incumbent cannibalises itself early”, we forced that move in all one hundred worlds and compared each world against its own baseline, same seed. The baseline, harvest first and pivot later, ended on a higher final position than the early pivot in 57 of 100 paired worlds. The early pivot won 39 and 4 tied. That is closer to even than a headline would make it, and the report puts the median cost of pivoting early at one point of final position. Pairing on the seed stops anyone mistaking a lucky run for a better strategy.
  • Fixed rules means reproducible, not certain. There is no ordinary randomness anywhere in the resolution path. Every roll comes from a seeded generator, a random-number source that gives the same sequence for the same seed, and the event that used the roll records its face. World number seventeen takes its seed from the run’s master seed and the number seventeen. Re-run it and you get the same world back, roll for roll. It is repeatable arithmetic, and it says nothing about the future.

Six things come out the other end

  • The beats. A timeline of what happened, beat by beat: the moves tried, the ones that failed, and why the arbiter priced them the way it did. Trying the same move again is a mechanic with a cost. In the live run the challenger tried the same public play five times, and the arbiter raised the target each time.
  • The dashboard series. The index lines beside the timeline are drawn by hand with the scenario, so they stay consistent with the runs. They are not measurements, and both the chart note and the report’s method section say so. They show the shape the author argues. Do not quote them as data.
  • The forks and the endings. “The Throat to Choke” ends on three endings, and each carries a named load-bearing assumption: the thing you have to believe for that ending to hold. Pick your ending and you will notice what you had to believe to pick it.
  • The evidence pack and its citation markers. Every bracketed marker like [s-003] in a report points to a claim in an evidence pack. Each claim carries a publisher, a date, a verbatim quote, a retrieval date and a verdict. A linter, a script that checks the prose, then holds the report to it. A sentence with a number in it has to cite a run event or a pack claim. Citing an id that is not in the pack is an error, because invented provenance is worse than none. Anything without a marker is authored, and the reports say that in plain words.
  • A seams section. Both reports carry one. It says where the machine surprised the author, where the author overrode it, and where the verified evidence cuts against the scenario. On “The Throat to Choke” that section lists eight places the evidence pushes back.
  • A named human writes the argument. Both published reports carry Dr Dan Epstein’s byline. The engine produces a record. A person makes the argument, the framing and the judgement calls, and that person is named on the page.

How much of each report a machine played

The answer differs per report, and it matters. So each report states it on its own face instead of leaving you to assume.

  • The Throat to Choke - human refereed. The author’s own account: an AI generated the five actor plays and Dr Dan Epstein refereed them. The report’s seams section then argues against its own material. One example: all five plays converge on “human judgement is the durable moat”, which is a suspiciously comfortable finding when an AI wrote the first pass.
  • The Seat Price - machine played. No human refereed this one while it ran. Four model seats played the archetypes in a five-beat live run, each with a private brief, win conditions and red lines. An AI arbiter set a difficulty for every claimed move and wrote the case for and against into the transcript. Dice decided. Separately, the fixed-rules version ran the full eight beats one hundred times. Dr Dan Epstein commissioned the question, set the table, read every transcript and wrote the account.
  • Where the human wrote past the run. In “The Seat Price”, the live run stopped at beat five. Beats six through eight are the author’s work, written on top of the hundred fixed-rules runs, and no model seat played them. So “machine played” is true of five beats in eight, and the ending is authored. The 2029 “dam break”, the beat where the built-up pressure on the secret of what each customer pays gives way, is the author’s reading of the game’s pressure mechanics. No run produced that beat as written. The report discloses this, and we repeat it here.

Whether the players are any good is a separate question

Everything above describes what the game does. It says nothing about whether an agent seated as a regulator behaves anything like a regulator. That is the question the whole exercise rests on. It has its own page, its own tests and its own failures.

The short version, and the only version this page may give: a reader can tell the characters apart, and their personalities change what they decide, not only how they write. Whether that makes them accurate about real behaviour is still an open question. Our first test of it was too blunt to read either way, and we withdrew the headline we had put on it. Nobody should claim accuracy in either direction, including us.

Read the validation page for what held, what broke, and the fault we found in our own product. The four passed studies each have a full technical report behind them:

One more limit. The machine can run this table one hundred times in an afternoon. It still cannot sit your leadership team in the seats and make them defend their assumptions to each other out loud. That is the tabletop, where your team plays the seats and the engine referees, and the difference is the one between reading a fire drill and running one.

The annex · Deep

The tables the prose describes: how stated odds become a target, and what an event carries.

How stated odds become a target number

A seat states the odds it thinks its move deserves. The engine converts those odds to a target number on a twenty-sided die. The arbiter then shifts the target by how the arguments for and against balanced out. This page reads the table below straight from the engine, so it cannot drift from the code.

Stated oddsTarget on a d20What the engine calls it
99 per cent1almost certain
95 per cent2highly likely
90 per cent3very likely
75 per cent6probable
50 per cent11even odds
25 per cent16probably not
10 per cent19highly unlikely
5 per cent20extremely unlikely
Below the table25very unlikely
Below that again30near impossible

The last two targets sit above what a twenty-sided die can reach on its own. A seat’s resource level adds between minus two and plus two to the roll, and the arbiter can itemise other modifiers on top of that. Without enough of them, a target of 25 or 30 is the arbiter’s way of ruling a move out while still rolling for it on the record. A natural 1, the die itself showing 1, always fails. A natural 20 clears any target of 20 or under.

What an event carries

The log is the run. Everything else in a report derives from it.

EventWhat it records
Every eventAn id, the tick (the beat number) it belongs to, a timestamp for audit, and the branch (which timeline) it sits on. The original timeline is called "main". A what-if replay gets its own label.
A claimThe seat, the claim in one declarative sentence, the odds that seat argues for it, the arguments for, and the arguments the other seats put against it.
A rulingThe decision in plain terms, the reasoning path, the assumptions the ruling rests on, and a self-reported confidence. Also the inputs weighed, which arbiter mode produced it, and whether it needed human sign-off.
A resolved checkThe target, the die face or faces, the itemised modifiers, the total, the outcome, and the margin by which it cleared or missed.
A reasoning journalWhat the agent considered, what it rejected and why, and how confident it was in the line it took.
A human overrideWho intervened, what they did, why, and which event they overrode. An override is an event like any other.

Re-run any world with its seed and you get it back, roll for roll. That is what reproducible means here, and it is all it means.