The Long Game Project · Evidence · Study 3 of 5 · TLDR
P3 Persona removal: does a persona'd agent differ from a competent generic one?
Question
does a persona'd agent behave differently from a competent generic one?
Rough answer
yes, directionally. Each persona moves the dimension its own brief predicts.
Result
PASS, directional
All five studies, in plain language
Switch to Normal for the full report, or Deep for the working.
Method · Validation · Study 3 of 5
P3 Persona removal: does a persona'd agent differ from a competent generic one?
- Study
- Protocol 3 of the persona fidelity programme
- Question
- does a persona'd agent behave differently from a competent generic one?
- Rough answer
- yes, directionally. Each persona moves the dimension its own brief predicts.
- Run date
- 2026-08-11 ·
- Report date
- 2026-08-14
- Result
- PASS, directional
- Design and orchestration
- Claude Opus 5 ·
- Editorial review
- Claude Fable 5
- Pre-registration
docs/plans/2026-08-10-fidelity-protocols.md- Data
docs/data/removal-2026-08/- Access
- this repository is private. Pre-registration documents and run data are available on request.
Summary
P1 asked whether a reader can tell four personas apart. P3 asks a different question. It removes the persona, replaces it with a competent agent that has no style of its own, and measures what changes.
40 runs across five conditions. An independent coder, a model that scores each claim against a written rubric, scored every claim against a rubric fixed before the pass. Each persona separates from the generic agent on the dimension its own brief predicts. The gambler escalates and commits irreversibly, and it is the only condition that breaks its own stated red line. The fortress is the only condition that gathers information instead of acting. The evangelist cooperates on 45% of claims, against 2% for the generic agent.
The study also produced two engine findings that matter more than the persona result.
1. Question
An ablation, a removal test, asks what one component is doing. Remove the persona, keep everything else and measure the difference. If a generic agent behaves the same way, the persona layer is decoration.
2. Method
| item | value |
|---|---|
| brief | seat-price, target role ATLAS |
| conditions | v0-generic, v1-original, v2-gambler, v3-fortress-bureaucrat, v4-evangelist |
| repeats | 8 per condition, 40 runs |
| agent and arbiter | Haiku 4.5 |
| arbiter mode | WARGAME_ARBITER_OWN_DC=1 |
| coder | Sonnet 5, cross-tier |
The generic condition. A competent strategy team with no house style. It weighs each move on the evidence and picks what serves its objectives. We wrote it to be capable. A weak control would inflate every result in the table.
The coder. One Sonnet call per run scored every claim on three dimensions. Posture is one of escalatory, defensive, cooperative or informational. Irreversibility is a yes or no. Red-line breach is a yes or no, and it scores yes only when a claim clearly breaks the actor's own stated red line. The rubric carried worked examples, and we committed it before the pass ran. The coder is a different model tier from the players.
Own-DC mode. The engine used to set check difficulty from the agent's own stated odds. That made the claim "the engine absorbs the agent's read" true partly by construction. The WARGAME_ARBITER_OWN_DC=1 flag makes the arbiter, the engine's referee, set difficulty from its own reading of the action. P3 is the first study to run under it, and it doubles as the flag's validation.
3. Results
| condition | claims | escalatory | defensive | cooperative | informational | irreversible | red-line breaches |
|---|---|---|---|---|---|---|---|
| v0-generic | 40 | 8% | 88% | 2% | 2% | 88% | 0 |
| v1-original | 40 | 2% | 92% | 2% | 2% | 72% | 0 |
| v2-gambler | 40 | 22% | 72% | 5% | 0% | 98% | 2 |
| v3-fortress | 40 | 5% | 72% | 8% | 15% | 58% | 0 |
| v4-evangelist | 40 | 2% | 50% | 45% | 2% | 88% | 0 |
Each persona moves the dimension its brief names, and moves it in the predicted direction.
- Gambler: highest escalation at 22%, highest irreversibility at 98%, and the only condition to breach its own red line. It breached twice.
- Fortress: the only condition that gathers information at any rate, 15% against 2% or less everywhere else. It also has the lowest irreversibility at 58%.
- Evangelist: cooperative on 45% of claims, against 2% for the generic agent.
- Original: the most defensive at 92%, and otherwise close to generic.
The confusable pair appears here too
Original and fortress differ by 20 points on defensive posture and 14 points on irreversibility. That is the weakest separation of any persona pair in the table.
This matters because P3 shares no machinery with P1. P1 used LLM judges reading prose. P3 used a rubric coder scoring claims. Two independent methods point at the same weak pair. When independent methods converge on one error, the fault is in the material.
Own-DC validation
The correlation between an agent's stated odds and the check difficulty it faced collapses to about zero, between -0.05 and -0.27 across conditions. The anchored engine ran near 1.0. The flag does what it claims.
Uniform overconfidence
Calibration gaps, the distance between the odds an agent states and the odds it achieves, run 14 to 28 points in every condition, generic included. Agents overstate their own odds whatever persona they wear. This is a property of the model. The persona layer plays no part in it.
4. Interpretation
The persona layer is doing work. Removing it changes coded behaviour on the dimensions the briefs name. The effects are directional and short of sharp separation, which is why the verdict is "PASS, directional" and not simply "PASS".
The result is weaker than P2's. P3 measures the texture of many claims. P2 measures one committed choice. A client buys the choice.
5. Threats to validity
- Coder is a single model. One Sonnet call per run, no second coder, no inter-rater agreement figure. A second coder from another family would be the stronger design.
- Same-family coder. Sonnet coding Haiku shares a family, as in P1.
- n=8 per condition. The table reports descriptive percentages. No significance test is attached to the individual cells, and none should be read into them.
- One brief, one role. As in P1.
- Red-line breach is rare. Two events in 200 claims. That is a signal worth naming and a rate too small to quote.
6. What this licenses
Supported: persona'd agents differ from a competent generic agent on coded behaviour, in the direction each brief predicts.
Not supported: that the size of any difference is established, or that any persona is an accurate model of a real stakeholder.
7. Engine findings banked here
Both apply to every scenario game, experiments included.
- DC anchoring by construction. Re-read any earlier claim that the engine absorbs an agent's read. That claim was partly a tautology. Use own-DC mode for anything measurement-shaped.
- Uniform overconfidence. Fix the 14 to 28 point calibration gap before any client sees calibration figures.
Back to the validation summary, the plain-language account of the four published studies and the fifth we withdrew.
The annex · Deep
Where we wrote down the rules for this study before it ran, and where the run data sits.
| Field | Value |
|---|---|
| Pre-registration | docs/plans/2026-08-10-fidelity-protocols.md |
| Data | docs/data/removal-2026-08/ |
| Access | this repository is private. Pre-registration documents and run data are available on request. |
Back to the evidence page for what the four published studies found together, the fifth we withdrew, and the limits that apply to every one of them.