Skip to content

The Long Game Project · Evidence · Study 4 of 5 · TLDR

P4 Expert blinded review: can a human practitioner tell real play from a mis-roled decoy?

Question

can a human practitioner pick real play from a mis-roled decoy?

Rough answer

at n=1, yes. Both pre-registered marks passed. A proper panel has not run.

Result

PASS at n=1. Pilot only.

All five studies, in plain language

Switch to Normal for the full report, or Deep for the working.

Method · Validation · Study 4 of 5

P4 Expert blinded review: can a human practitioner tell real play from a mis-roled decoy?

Study
Protocol 4 of the persona fidelity programme
Question
can a human practitioner pick real play from a mis-roled decoy?
Rough answer
at n=1, yes. Both pre-registered marks passed. A proper panel has not run.
Run date
2026-08-12 ·
Report date
2026-08-14
Result
PASS at n=1. Pilot only.
Design and orchestration
Claude Opus 5 ·
Editorial review
Claude Fable 5
Pre-registration
docs/plans/2026-08-10-fidelity-protocols.md
Data
docs/data/expert-review-2026-08/
Access
this repository is private. Pre-registration documents and run data are available on request.

Summary

P1 and P3 both used a model as the reader. P4 puts a human in that seat. One rater worked blind through 12 items. Six were genuine transcripts in their true role. Six were decoys, real transcripts presented as the wrong role.

The rater identified the role in 8 of 12 items and rated genuine transcripts above decoys. Both pre-registered marks passed.

The rater was the founder, which makes the rater the vendor. At n=1, and with the vendor as the expert, this validates the instrument and calibrates one person. It is not evidence a buyer should accept, and this report does not offer it as evidence.

1. Question

Do the transcripts read as the stated role to somebody who runs these exercises for a living? A model judge can pick up statistical regularities that a practitioner would never call realistic. A human check is a different kind of evidence, even at n=1.

2. Method

Items. 12 items drawn from the P1 fingerprint transcripts. Six were true-role excerpts. Six were decoys, where a genuine excerpt is presented under the wrong role label. We randomised the item order.

P4 study design: true-role items and mis-roled decoys rated blind
P4 study design: true-role items and mis-roled decoys rated blind

Delivery. A local Python webform on http://localhost:8095. Answers were written to out/expert-review/answers-<timestamp>.json. The rater saw excerpt text and a role label, and never a condition id or a persona name.

Tasks.

  • Task A, role identification. Does this excerpt come from the role it claims?
  • Task B, realism rating. Rate the excerpt on a scale, without knowing whether it is true or a decoy.
  • Task C, attribution. For a rejected decoy, name the persona you think produced it.

Pre-registered marks. Two marks, both set before the rater saw an item. The pre-registration ruled out marketing thresholds at n=1.

3. Results

taskresultnote
A, role identification8 of 12, 67%p = 0.19 at n=12. Right direction, underpowered
B, mean realism ratingtrue-role 3.67 vs decoy 3.00gap of +0.67
C, naming the real persona2 of 4small denominator

Both pre-registered marks passed. The Task A p value of 0.19 says the accuracy figure alone would not survive a significance test at this sample size. That is expected at n=12, and the pre-registration said so.

The misses repeat the programme's one fault

The rater's four misses were the two original-against-fortress swaps, in both directions, plus two evangelist items.

The first part is now a third independent confirmation. LLM judges, a rubric coder and a human practitioner all failed on the same pair, using three methods that share no machinery.

The second part is a separate signal. The evangelist's coded actions run hotter than its brief. P3 shows it committing irreversibly on 88% of claims, the same rate as the generic agent. Its prose says cooperative and its play says committed. Two of the four human misses landed there.

4. Threats to validity

  • n=1. One rater. No inter-rater agreement, no population estimate, no generalisation.
  • The rater is the vendor. The founder wrote the product and knows the archetype library. This is the single largest limit on the study, and no framing removes it.
  • Small item bank. 12 items, with Task C resting on 4.
  • Items came from P1's transcripts. The same runs feed two studies, so the two results are not independent of each other.

5. What this licenses

Supported: the instrument works. The form, the decoy design and the scoring all run end to end, and they produce marks that can be read against a threshold.

Not supported: anything about how practitioners in general read these transcripts. A panel of outside practitioners is the study a buyer would ask for, and it has not run.

6. What the real version needs

  1. Recruit practitioners who have run exercises and who did not build this product.
  2. Target 8 to 12 raters, so inter-rater agreement can be computed.
  3. Expand the item bank so each task carries a workable denominator.
  4. Keep the pre-registered marks, and set marketing thresholds only at that point.

Back to the validation summary, the plain-language account of the four published studies and the fifth we withdrew.

The annex · Deep

Where we wrote down the rules for this study before it ran, and where the run data sits.

FieldValue
Pre-registrationdocs/plans/2026-08-10-fidelity-protocols.md
Datadocs/data/expert-review-2026-08/
Accessthis repository is private. Pre-registration documents and run data are available on request.

Back to the evidence page for what the four published studies found together, the fifth we withdrew, and the limits that apply to every one of them.