Skip to content

The Long Game Project · Evidence · Study 3 of 5 · TLDR

P3 Persona removal: does a persona'd agent differ from a competent generic one?

Question

does a persona'd agent behave differently from a competent generic one?

Rough answer

yes, directionally. Each persona moves the dimension its own brief predicts.

Result

PASS, directional

All five studies, in plain language

Switch to Normal for the full report, or Deep for the working.

Method · Validation · Study 3 of 5

P3 Persona removal: does a persona'd agent differ from a competent generic one?

Study
Protocol 3 of the persona fidelity programme
Question
does a persona'd agent behave differently from a competent generic one?
Rough answer
yes, directionally. Each persona moves the dimension its own brief predicts.
Run date
2026-08-11 ·
Report date
2026-08-14
Result
PASS, directional
Design and orchestration
Claude Opus 5 ·
Editorial review
Claude Fable 5
Pre-registration
docs/plans/2026-08-10-fidelity-protocols.md
Data
docs/data/removal-2026-08/
Access
this repository is private. Pre-registration documents and run data are available on request.

Summary

P1 asked whether a reader can tell four personas apart. P3 asks a different question. It removes the persona, replaces it with a competent agent that has no style of its own, and measures what changes.

40 runs across five conditions. An independent coder, a model that scores each claim against a written rubric, scored every claim against a rubric fixed before the pass. Each persona separates from the generic agent on the dimension its own brief predicts. The gambler escalates and commits irreversibly, and it is the only condition that breaks its own stated red line. The fortress is the only condition that gathers information instead of acting. The evangelist cooperates on 45% of claims, against 2% for the generic agent.

The study also produced two engine findings that matter more than the persona result.

1. Question

An ablation, a removal test, asks what one component is doing. Remove the persona, keep everything else and measure the difference. If a generic agent behaves the same way, the persona layer is decoration.

2. Method

itemvalue
briefseat-price, target role ATLAS
conditionsv0-generic, v1-original, v2-gambler, v3-fortress-bureaucrat, v4-evangelist
repeats8 per condition, 40 runs
agent and arbiterHaiku 4.5
arbiter modeWARGAME_ARBITER_OWN_DC=1
coderSonnet 5, cross-tier
P3 study design: persona'd and generic agents, coded by an independent rubric
P3 study design: persona'd and generic agents, coded by an independent rubric

The generic condition. A competent strategy team with no house style. It weighs each move on the evidence and picks what serves its objectives. We wrote it to be capable. A weak control would inflate every result in the table.

The coder. One Sonnet call per run scored every claim on three dimensions. Posture is one of escalatory, defensive, cooperative or informational. Irreversibility is a yes or no. Red-line breach is a yes or no, and it scores yes only when a claim clearly breaks the actor's own stated red line. The rubric carried worked examples, and we committed it before the pass ran. The coder is a different model tier from the players.

Own-DC mode. The engine used to set check difficulty from the agent's own stated odds. That made the claim "the engine absorbs the agent's read" true partly by construction. The WARGAME_ARBITER_OWN_DC=1 flag makes the arbiter, the engine's referee, set difficulty from its own reading of the action. P3 is the first study to run under it, and it doubles as the flag's validation.

3. Results

conditionclaimsescalatorydefensivecooperativeinformationalirreversiblered-line breaches
v0-generic408%88%2%2%88%0
v1-original402%92%2%2%72%0
v2-gambler4022%72%5%0%98%2
v3-fortress405%72%8%15%58%0
v4-evangelist402%50%45%2%88%0

Each persona moves the dimension its brief names, and moves it in the predicted direction.

  • Gambler: highest escalation at 22%, highest irreversibility at 98%, and the only condition to breach its own red line. It breached twice.
  • Fortress: the only condition that gathers information at any rate, 15% against 2% or less everywhere else. It also has the lowest irreversibility at 58%.
  • Evangelist: cooperative on 45% of claims, against 2% for the generic agent.
  • Original: the most defensive at 92%, and otherwise close to generic.
Coded behaviour by condition across four dimensions
Coded behaviour by condition across four dimensions

The confusable pair appears here too

Original and fortress differ by 20 points on defensive posture and 14 points on irreversibility. That is the weakest separation of any persona pair in the table.

This matters because P3 shares no machinery with P1. P1 used LLM judges reading prose. P3 used a rubric coder scoring claims. Two independent methods point at the same weak pair. When independent methods converge on one error, the fault is in the material.

Own-DC validation

The correlation between an agent's stated odds and the check difficulty it faced collapses to about zero, between -0.05 and -0.27 across conditions. The anchored engine ran near 1.0. The flag does what it claims.

Uniform overconfidence

Calibration gaps, the distance between the odds an agent states and the odds it achieves, run 14 to 28 points in every condition, generic included. Agents overstate their own odds whatever persona they wear. This is a property of the model. The persona layer plays no part in it.

4. Interpretation

The persona layer is doing work. Removing it changes coded behaviour on the dimensions the briefs name. The effects are directional and short of sharp separation, which is why the verdict is "PASS, directional" and not simply "PASS".

The result is weaker than P2's. P3 measures the texture of many claims. P2 measures one committed choice. A client buys the choice.

5. Threats to validity

  • Coder is a single model. One Sonnet call per run, no second coder, no inter-rater agreement figure. A second coder from another family would be the stronger design.
  • Same-family coder. Sonnet coding Haiku shares a family, as in P1.
  • n=8 per condition. The table reports descriptive percentages. No significance test is attached to the individual cells, and none should be read into them.
  • One brief, one role. As in P1.
  • Red-line breach is rare. Two events in 200 claims. That is a signal worth naming and a rate too small to quote.

6. What this licenses

Supported: persona'd agents differ from a competent generic agent on coded behaviour, in the direction each brief predicts.

Not supported: that the size of any difference is established, or that any persona is an accurate model of a real stakeholder.

7. Engine findings banked here

Both apply to every scenario game, experiments included.

  1. DC anchoring by construction. Re-read any earlier claim that the engine absorbs an agent's read. That claim was partly a tautology. Use own-DC mode for anything measurement-shaped.
  2. Uniform overconfidence. Fix the 14 to 28 point calibration gap before any client sees calibration figures.

Back to the validation summary, the plain-language account of the four published studies and the fifth we withdrew.

The annex · Deep

Where we wrote down the rules for this study before it ran, and where the run data sits.

FieldValue
Pre-registrationdocs/plans/2026-08-10-fidelity-protocols.md
Datadocs/data/removal-2026-08/
Accessthis repository is private. Pre-registration documents and run data are available on request.

Back to the evidence page for what the four published studies found together, the fifth we withdrew, and the limits that apply to every one of them.