Skip to content

The Long Game Project · How it works · Evidence · TLDR

What we can and cannot claim

The characters are distinct and they change what the AI players, the agents, decide. Whether they are accurate is still open. We found a fault in our own product and fixed it.

Blinded readers naming the characterBlinded readers named the character behind an anonymised transcript 79 times in 100, against a 25 in 100 guessing baseline.Readers namedthe character79 %Chance25 %0100 %
Blinded readers naming the characterBlinded readers named the character behind an anonymised transcript 79 times in 100, against a 25 in 100 guessing baseline.Readers named the character79 %Chance25 %0100 %
Blinded readers named the character behind an anonymised transcript 79 times in 100. Guessing would get 25. Study 1.

Can you tell the characters apart?

Yes, clearly.

Do they change what agents decide?

Yes.

Does a personality make an agent behave like the real organisation?

Unknown. Our first test could not read it either way.

Does the diagnostic we ship drive that behaviour?

Yes, now. Our first test said no, and fixing why made the product better.

Switch to Normal for the full page, or Deep for the working.

The Long Game Project · Method · Validation

What we tested, what held, what broke, and what we still cannot say

Every number on this page comes from a test whose rules we wrote down before it ran. No real organisation and no human respondent appears in any of those tests. Every figure is a model reading a written brief. The failures sit next to the passes, and one failure is a headline we took back ourselves. The retractions are the reason the surviving claims are worth anything.

New here? This page tests the players. If you do not yet know the game, who sits at the table, how a turn works and what comes out the other end, start with what the wargame engine is, then come back.

The instrument

The diagnostic asks 54 questions and scores an organisation on 15 strategic dimensions, the Ingredients as the rest of this site calls them. In August 2026 we ran a measurement programme against it: eight purpose-built organisation profiles, eight independent readings of each, and the rules for every run written down before we spent it.

What held. The instrument separates organisations reliably. Its mean ICC is 0.84 across the 15 dimensions. ICC, the intraclass correlation, is the share of score variation that is real difference between organisations rather than noise, so 0.84 means most of what the diagnostic reports is real. All fifteen dimensions read their own trait, with a mean correlation of 0.88 between what each was designed to measure and what it measured, and no two dimensions share a question. The weakest dimension, breadth of scope, gives a direction and no more. Read it that way.

What we retracted. Five interim claims we made during the programme were wrong, and we withdrew them. One was an early claim that five dimensions measured nothing. We had judged each dimension on test cases that held it constant, so the zero result came from our test design, not from the dimensions. Four of those five dimensions later scored ICC 0.75 to 0.96. The fifth did not reach that band, and none of the five was dead.

The ceiling. These numbers measure a model reading a written brief, which is a very consistent respondent. A human answering the same questions would be less consistent, so every figure here is a best case and no estimate of human test-retest reliability. Scores are not precise beyond about 8 points on a 100-point scale. Ask the instrument about the same organisation five times and it lands on the same overall character on four of the five passes. Compare every pair of passes and the agreement is 60 per cent. That consistency caps how right any reading built on the instrument can be, and we would rather measure that ceiling than assume it.

The players, in short

Our wargames seat AI agents as the other parties: the regulator, the competitor, the board. The whole exercise rests on one claim, that the agent playing the regulator behaves like a regulator. Until August 2026 nobody had tested that, including us. We ran five studies. Each one raised a check the next had to pass, and that kept us testing for another week. Four questions, four answers.

  • Can you tell the characters apart? Yes, clearly.
  • Do their personalities change what they decide, as well as how they talk? Yes.
  • Does giving an agent a personality make it behave more like the real person? Still unknown. Our first test could not tell either way.
  • Does the diagnostic we ship actually drive that behaviour? Yes, now. Our first test said no, and fixing why made the product better.

The first two are worth something. The third is the one people assume, and it is still open. The fourth is the fault we are most glad we found.

What held

You can tell the characters apart. We took transcripts of the AI playing, stripped every name out and gave them to blinded readers: AI models from a different tier than the players, meaning a bigger model from the same maker, which had never seen the character material. Those readers named the right character 79% of the time, against a 25% guessing rate. A result that strong would come up by chance about once in 25 million tries. A separate scoring system agreed. So did a human expert working through a blind test, but that expert was our own founder, one rater, so that pass checks the test design and nothing more.

The characters change what they decide, as well as how they talk. We gave the agents a genuine business fork with two defensible answers: sign a binding three-year exclusive, or take a shorter pilot you can walk away from, which costs a bit more later. We wrote down beforehand which character should pick which. The commercially-minded character, the merchant in the study reports, signed the binding deal 9 times out of 9 (one of its ten runs never reached the fork). The cautious, process-bound character, the fortress, refused it 10 times out of 10 and took the pilot. No overlap. A split that clean would come up by chance about once in 90,000 tries.

That fork exists because of an earlier fault. Every one of our first checks made the same mistake: none could tell our standard incumbent character from our cautious bureaucrat. When three independent readers make the same error, the fault sits in the material. The cause was simple. The two characters differed only by degree, one cautious and the other more cautious. We rewrote both so they differ in kind, then built a test to tell them apart. The pair nobody could separate is now the cleanest separation we have.

The accuracy test, and the headline we withdrew

Distinct and accurate are different things. An agent can be a recognisable cautious bureaucrat and still behave like no real bureaucrat who ever lived. Consistent is not the same as right.

So we ran the test. We rebuilt a documented situation, Blockbuster in mid-2004 facing Netflix, using only what was public at the time. We scored a researched character against a plain competent agent on reproducing the five moves Blockbuster made. Twenty runs, ten each way. They scored the same. For three days we reported that as a clean negative: the character made no difference.

Then we checked the test itself and withdrew that reading. Four of the five scored decisions came out the same way in almost every run, whoever was playing. Blockbuster against Netflix is also one of the most retold business stories there is, so the model already knew the ending. A test like that can only say “we could not see anything”. It cannot say “there is nothing to see”.

The accuracy question is still open, in both directions. We do not claim the characters make agents accurate, and we do not claim they fail to. We are building a sharper test: a case the model cannot already know, scored against decisions it cannot have memorised. We wrote down in advance that a null was possible and what it would mean. We did not anticipate a test too blunt to measure anything either way. Both facts are on the record.

The fault we found in our own product

Partway through, one check caught every study above: none of them had tested the product we ship. Every result so far used characters we wrote by hand. The product builds its characters from the 54-question diagnostic, and that pipeline had never been inside an experiment.

So we ran the missing test. Same business fork, same rules, with characters built by the diagnostic instead of by hand. They failed completely. A diagnostic-built character behaved like an agent with no character at all, while the hand-written pair kept its perfect split in the same session.

The diagnostic itself was sound. It measured the two organisations as clearly different, on the traits that decision turns on. The fault sat one step later. The code that turns those measurements into the character’s playing instructions was throwing the measurements away and using a stock template instead.

We fixed that step, changed nothing else and re-ran the same test. In the runs we made, at most ten per character, the diagnostic-built characters now split 90% against 10% on the same decision. If we had shipped without that test, every claim above would have been true of our demo material and false of the product. That is why the failures are on this page.

The studies themselves

Everything above is the short account. These are the studies behind it, one page each, with the method, the pass marks written down before the run, the raw counts and the threats to validity. They are technical by design. For what the exercise around these players does, see what the wargame engine is.

The fifth study, on historical accuracy, is not published. We withdrew its verdict when we found the test could not discriminate, and we will rewrite the report from the replacement study. What it supports today is the account above: the accuracy question is open, in both directions.

What this licenses, in plain terms

Fair to say: the diagnostic separates organisations reliably, each dimension measures its own trait, the wargame characters are distinct and recognisable, their personalities change their decisions by a measured margin, the diagnostic we ship now drives those decisions, and we test all of it.

Not fair to say: that the agents predict what a real regulator or competitor would do. We tried to test that, the test turned out to be too blunt to answer, and we are building a sharper one. Until it runs, nobody should claim accuracy in either direction, including us.

What that means in practice. These exercises are worth running to pressure-test your own thinking against other players who behave consistently and differently from each other. That is what a good tabletop exercise has always been for. They do not forecast, and our published reports carry the same rule: scenarios, not forecasts.

Two honest marks against the good result: we predicted twelve character choices in advance, four characters on three forks, and got ten right, with both misses on the two secondary forks rather than the main one. And the sample is at most ten runs per character. That is enough to show that two characters differ from each other. It is not enough to prove that each one differs from a plain agent, and we do not claim it does.

Other limits worth knowing. All the persona testing used one cheap AI model, in one kind of business scenario, in enterprise software, so none of it carries over to healthcare, finance or government on its own. The human expert in the blind test was our own founder, which validates the test design and nothing more. A panel of independent practitioners is the study a serious buyer should ask for, and it has not run yet.

Every agent, with a personality or without one, states odds for its moves 14 to 28 percentage points higher than the share of those moves it then wins. That overconfidence is a property of the underlying model, and we are fixing it. Every figure on this page comes from a model reading a written brief, so treat each one as a best case and no estimate. The four study reports carry the methods and the raw counts. Pre-registration documents and run data are available on request through the Long Game Project contact page.

The annex · Deep

The studies as a table: what each one asked, and the verdict its own report carries.

StudyQuestionResult
P1 Fingerprint recoverycan a blinded reader tell which persona produced a transcript, from play alone?PASS (strong)
P2 Contested forksdo personas change what an agent picks at a genuine fork, not just how it writes?v2 instrument FAILED. v3 instrument rebuilt and PASSES all gates.
P3 Persona removaldoes a persona'd agent behave differently from a competent generic one?PASS, directional
P4 Expert blinded reviewcan a human practitioner pick real play from a mis-roled decoy?PASS at n=1. Pilot only.
P5 Historical anchoringDoes a researched persona reproduce a documented decision?Withdrawn: the test could not discriminate. 20 runs.

We read every question and verdict above from the study report itself at build time, so this table cannot drift from the reports it summarises. P5 is not published, for the reason its row states.

The terms in the Result column are the study reports’ own. v2 and v3 are the second and third versions of the fork test and its character briefs: v2 failed to tell two characters apart, and v3 is the rewrite described on this page. Gates are the pass marks written down before a run. Directional means the effect points the way the brief predicted but stops short of a sharp separation.