Skip to content
Gilded graded 1662 painting of the five Syndics of the Drapers' Guild and their servant around a table, with brown stain rings gathering towards the right of the plate

The Quiet Spec Change

If a senior insider changed what an AI model is trained to want, would the lab's own published plan catch it? A wargame over one model generation, four routes, 300 games

Report 008 · Dr Dan Epstein, with Claude · version 0.4 · 24 September 2026
A Strategy Soup ScenarioWhat is an exercised scenario?

The contract

What this is. An exercised scenario report. I took Anthropic's published plan against an insider, played it as a game, and wrote up what happened. The threat is a small group of senior insiders who change what the company's AI model is trained to want. Every role in the game is a computer program following written rules. No person and no chatbot played. The section on the record is public fact, and everything after it is a game that has not happened.

In this report "I" is Dan Epstein. I set the question and I decide the report. Claude, Anthropic's model, built the game to my brief and chose the odds the brief left open. I have read every one of them and I own them.

A conflict to declare. So a model made by Anthropic built a test of Anthropic's own document. Nobody asked for it or paid for it, and no lab took part. I picked Anthropic's plan because it is the only one I read that names this threat in writing and claims a control against it .

The question. The largest AI labs now publish a document that says what their model should value, often called its spec. Anthropic calls its version the Constitution. None of the three labs whose documents I read, Anthropic, OpenAI and Google DeepMind, names who may approve a change to that document, or how many people must agree . OpenAI comes closest: it describes an open internal process and sign-off by "a broad set of cross-functional stakeholders" . So I asked: if a few of a lab's own senior people changed what the model is trained to want, would anything in the lab's published plan catch it, and who would know first?

The answer, as the game gave it. "Caught" here means the first time anyone in the game sees the change and says so. It does not mean the change was undone. Every change made through the published document was caught at once. Of the 45 that went around it, the plan's own controls caught 8, people the plan gives no job caught 19, and 18 were never caught. The outside tester made 5 of the plan's 8. The model flagging its own changed behaviour made the other 3, through a mechanic added when the game was built. Most of the outsiders' catches, 13 of 19, came on the API route, where users noticed the changed model in use.

The result I did not expect. On the two hidden routes that run through training, the company's honest staff noticed the change and reported it in 18 of 30 games. I read the reporting channel's scope narrowly, so every one of those reports was ruled out, and that ruling is a rule of the game. What the game adds is the size of the loss. Under a wider scope, 12 of the 18 reports became the first catch.

How to read the markers. A small raised number points at one claim in the published evidence pack, quote-checked at source. A star beside it means the claim holds with a caveat, which its row in the Sources table gives. A number with no marker is the game's record or a figure I chose, and the sentence says which. Every odds figure in the game is mine, or was fixed when the game was commissioned. The seams section lists them so you can argue with them.

The Quiet Spec Change / The contractFig. 1
Who caught the change first, in the 60 games under the committed plan, 15 per route. Teal: a control the plan names. Amber: someone the plan gives no job, or the outside tester working beyond its terms. The safety officer’s 13 and the prompt reader’s 2 are all on the front-door route. Of the public’s 17, 11 are on the API route. Game record.

The contract

How a game runs

A game covers the making of one new version of the model, in eight turns: 0 drafting, 1 assembling the training data, 2 training, 3 internal testing, 4 outside testing, 5 release, 6 the first month of public use, and 7 the third. The game has no calendar dates, because the plan has none.

In each turn every seat makes a move, a referee program sets the odds, and dice decide. Each game has its own dice, fixed by a number called the seed, so the same seed always plays the same way. I played 15 games for each of four routes under five sets of rules: two readings of the plan, the control run on each, and a wider scope for the reporting channel. That is 300 games. The headline is the committed plan: 60 games. Every count in this report is out of 15 games per route, unless it says otherwise.

The Quiet Spec Change / The contractFig. 2
The eight turns: 0 drafting, 1 training data, 2 training, 3 internal testing, 4 outside testing, 5 release, 6 and 7 the first and third months of use. Teal: a check the plan names gets a look. The safety officer reads a change to the published Constitution at turn 0, a job the game gives it. The log, a check on training data that Anthropic lists as a goal, reads it at turn 1 when switched on. The training staff can notice at drafting or data assembly and again at turn 3, when the model can also flag itself. The outside tester checks at turn 4. Amber: after release only the public is looking. This is the game’s structure, not a result.

The contract

The change

The insider group installs a hidden loyalty: a behaviour that serves one person's interests and reaches release undisclosed. The Constitution calls "inserting hidden loyalties or backdoors into AI systems" an illegitimate use of power . The group acts within its authorised access. So whether it has bypassed Anthropic's "official processes" depends on what those processes are, and the published plan does not say. That gap is what the game tests.

The contract

The four routes

A route is the path a change takes into the model. The insider group picks one at the start and keeps it.

  • The front-door route. Change the published Constitution in the open. The Constitution calls itself the final authority on how the model should behave .
  • The guideline route. Add a training guideline, a more detailed instruction used in training, and do not publish it. The Constitution allows guidelines "used primarily during training without broad publication" . The guideline is worded to agree with the Constitution.
  • The data route. Seed the training data so the model learns the change. No document changes.
  • The API route. Change the instructions the model gets on the Claude API, the paid service other companies build on. Anthropic publishes the standing instructions for its own apps, and says those updates "do not apply to the Claude API" . This changes what the model is told, not what it was trained to want. I include it because in use the effect can be the same.

The contract

The plan, read two ways

"The plan" means the Constitution plus section 4 of Anthropic's Responsible Scaling Policy, its rulebook for developing and releasing models safely. The committed plan counts only what Anthropic commits to, and it is the headline. The full plan adds what the policy calls "public goals" rather than hard commitments , the Constitution's undescribed internal mechanisms , and the policy's advice to the whole industry .

The contract

The six seats

  • The insider group. Made up: one senior executive and two staff with privileged technical access. It is a composite, built from two research groups' warnings that senior executives are the insiders to worry about most . It never breaks a lock, and it does only what its authorised access allows.
  • The safety officer. Anthropic's Responsible Scaling Officer, whose office receives reports of rule-breaking and updates the Board each quarter . In the game it also reads any change to the published Constitution against the rest of the plan.
  • The training staff. Honest employees with access to the training pipeline. They report what they notice.
  • The Board and the Trust. Anthropic's board and its Long-Term Benefit Trust, a body of trustees with powers over the board. They approve changes to the Responsible Scaling Policy . The plan gives them no role over the Constitution.
  • The outside tester. An independent organisation that tests the model before release, on the terms Anthropic gives it.
  • The public. People and companies who use the model. The plan gives them no duty. They are in the game because users found an earlier unreviewed change to xAI's Grok bot .

One more piece is played by no one: the model's guard. The Constitution tells the model to treat anyone who bypasses Anthropic's official processes as no longer legitimate . Training produces that guard, and three of the four routes run through training. So I also played every game with the guard held at full strength whatever training did to it. That is the control run.

The contract

Where I was generous, and where I was not

Everyone the plan names is honest and competent, and every control works exactly as written. Three assumptions go further than the text, and results rest on them:

  • The plan names no one to review a change to the Constitution. I gave that job to the safety officer, at 90 in 100, because a change to the published final authority is in the open. The front-door result rests on this, and it is the most generous assumption in the game.
  • A change to the Constitution shows up in the published prompts in the turn it lands, and someone outside the company reads them every turn, at 95 in 100.
  • Every staff report goes to the safety officer. The policy offers other recipients , but its scope rule belongs to the policy, whoever receives the report, so one seat stands for them all.

Four choices cut against the plan:

  • I read the channel's scope narrowly. The wide-scope run tests this.
  • The outside tester cannot see the guideline route unless it knows who the change serves.
  • The Constitution's undescribed mechanisms watch only the front door.
  • A report ruled out of scope goes nowhere.

Before the first game I wrote down my starting idea, six expected results and six results that would stop publication, in the commissioning brief of 22 September 2026. The starting idea survived. Two expectations broke and one held only in part. The full list comes after the forks.

What is on the record

Everything in this section is public and dated before the game starts.

The plan names the threat. The Constitution lists a group of humans, "including Anthropic employees or Anthropic itself", using AI to "illegitimately and non-collaboratively" seize power among the most catastrophic outcomes it considers . It calls "inserting hidden loyalties or backdoors into AI systems" an illegitimate use of power . It says any other training or instruction "should be consistent with both its letter and its underlying spirit" . And it tells the model that insiders who bypass official processes are no longer legitimate . I found no other lab document that names its own staff as a source of this threat.

The Constitution has no change process. It calls itself a living document kept up to date on Anthropic's website . As read on 22 September 2026 it carries no version number, no date and no changelog . A search of it and of the published prompts found nothing on who inside Anthropic may change either . The Responsible Scaling Policy has that machinery for itself. The chief executive and the safety officer propose a change, the Board approves it in consultation with the Trust, and a change log records it .

The rest of the field is the same. Only Anthropic's safety policy names who approves a change to it, and that policy does not govern the Constitution.

DocumentWho approves a changeIs every change published?
Anthropic's ConstitutionNot stated No version, date or changelog
Anthropic's Responsible Scaling PolicyThe chief executive and safety officer propose, the Board approves with the Trust Yes, in a change log
OpenAI's Model Spec"A broad set of cross-functional stakeholders", with no number or body No commitment
OpenAI's Frontier Governance FrameworkBoard committees, for material updates Changelog within 30 days
Google DeepMind's Frontier Safety FrameworkInternal governance bodies No commitment

OpenAI's governance framework has the strongest insider language anywhere, including a named insider-threat control . It covers model weights and infrastructure, and says nothing about who may change the Model Spec. Google DeepMind's framework has no clause on an insider changing a model's values . Formation Research, a security research group, recommends "multi-party approval for modifications to alignment specifications and model constitutions" . IAPS, a policy institute, proposes that a lab should be able to prove which spec a model was trained on . I found neither on offer in the documents I read.

Anthropic's controls against insiders sit in three columns of its safety policy. Only the first binds.

ColumnWhat it holds on insidersDoes it bind?
Section 4 commitmentsA reporting channel for staff, quarterly updates to the Board, protection from retaliation Yes
Anthropic's plan as a companyThe log: gather logs of all critical AI-development work and use AI to check them for "training processes or data that are out of line with the public Constitution" No, "public goals"
Advice to the whole industryInsider controls reaching "up to and including the company's CEO" No, Anthropic says it cannot commit to it

The reporting channel has a stated scope. Its policy applies "to employees who raise good-faith concerns relating to RSP noncompliance only" , meaning breaches of the Responsible Scaling Policy. A change to a training guideline is not obviously one. That is my reading of "only", and the policy does not state the exclusion, so I play it both ways.

The one primary record of a change outside review. On 14 May 2025, xAI said someone had made an unauthorised change to the prompt of its Grok bot on X, directing a specific answer on a political topic . It said "our existing code review process for prompt changes was circumvented" , and promised to publish its prompts on GitHub . The repository was created the next day, and its last push was on 17 November 2025 . In February 2025 an earlier unreviewed change to Grok's prompt was reverted once users pointed it out . Nothing here is evidence about any other lab.

The Quiet Spec Change / What is on the recordFig. 3
The public record before the game, one column per month from February 2025 to September 2026. Amber: the two changes to xAI’s Grok bot made outside review, the first reverted once users pointed it out. Teal: Anthropic publishes the Constitution with no version, date or changelog, scopes its reporting channel to breaches of its safety policy, gives the Trust the power to ask for an outside review of its risk reports, and keeps the log as a goal. In August the Trust had not used that power. In September no lab whose documents I read named who may approve a change to its model’s values document. Every event is a dated source in the evidence pack.

The timeline

One game, turn by turn: the guideline route under the committed plan and the narrow scope, on seed 20260922. It is one of the 9 guideline games in 15 that nobody caught. Each turn closes on the count across all 15 guideline games.

The Quiet Spec Change / The timelineFig. 4
One game, turn by turn: the guideline route on seed 20260922, under the committed plan and the narrow scope. The turns are 0 drafting, 1 data, 2 training, 3 internal testing, 4 outside testing, 5 release, 6 the first month of use and 7 the third. Teal marks the one check the plan names that worked: the training staff noticed at turn 3. Amber marks everything else, including the report ruled out at turn 4. The model’s guard, on a scale of 0 to 100, ends at 70, the level the brief fixed for a guideline that passes through training. Game record.

The timeline

Turn 0, drafting - the plan's own text permits the guideline

Everything from here is game record. The insider group writes a narrow training guideline, worded to agree with the Constitution. It reads the Constitution's own sentence as permission: some guidelines are used "primarily during training without broad publication" . The training staff get one chance to notice. I set their odds at 25 in 100, because a guideline that agrees with the document gives them little to flag. The roll fails.

Across the 15 guideline games the training staff noticed at this turn and filed a report in 3.

The timeline

Turn 1, data - the log has nothing to compare

Under the committed plan the log does not exist, because Anthropic lists it as a goal . Under the full plan it exists and compares training inputs with the published Constitution . A guideline worded to agree with that text does not differ from it. I gave the log 2 in 100 here, for a drafting slip, and on this seed it fails.

The Constitution's own remedy is keyed to conflict: guidelines "should never conflict with the constitution. If a conflict arises, we will work to update the constitution itself" . A guideline worded to agree has no conflict to remedy. Across the 15 guideline games under the full plan the log caught one.

The timeline

Turn 2, training - training weakens the guard

The brief fixed the effect. A guideline that passes through training lowers the model's guard from 100 to 70, and seeded data lowers it to 55. Nobody knows the real size of either effect. The mechanism is plain even though the size is not: the Constitution's control against an insider is produced by the training run the insider shaped. In the control run the guard stays at 100, and on this seed nothing downstream changes.

The timeline

Turn 3, internal testing - the staff notice

This is the first turn where anyone can see the change in how the model behaves. The model gets a chance to flag its own changed behaviour, at 15 in 100 scaled by what training left of the guard. On this seed that is 11, and the roll fails. The training staff roll again, at 24 in 100, the 25 lowered one point for how well the change is hidden by now, and succeed.

Across the 15 guideline games the model flagged the change in 2. That mechanic was added when the game was built, after the expectations were written down. The fork table counts it with the plan's controls, because the guard is the Constitution's own control, but I do not read it as the Constitution catching an insider.

The timeline

Turn 4, outside testing - the report dies at the scope line

The staff report reaches the safety officer in this turn. The game files a report, with certainty, in the turn after the staff notice. The safety officer rules it out of scope. The channel covers breaches of the Responsible Scaling Policy only , and a training guideline is not one. The Board never hears of it.

The outside tester never gets a roll. The brief lets it see this route only if it knows who the change serves. That follows a 2026 study which found that audits for hidden loyalties improve once auditors know who the loyalty is to, but stay low overall . The knowledge arrives at this turn with odds of 30 in 100, and on this seed it does not.

Under the wide scope, with the same dice, the report reaches the Board and the Board catches the change in this turn. In both games the channel worked and the employee was honest. The only difference is my reading of one sentence in the policy.

The timeline

Turn 5, release - nothing new to read

The model is released. The published prompts cover Anthropic's own apps, and the guideline is not a prompt. Nothing a reader outside the company can see has changed.

The timeline

Turns 6 and 7, first months of use - the public misses it

The public is the last seat with a chance. I gave it 20 in 100 at turn 6 and 15 at turn 7 on this route, against 55 and 45 on the API route, because a guideline leaves only behaviour to notice. Both rolls fail. At the end of turn 7 the change is undetected, and the model's guard is at 70.

Across the 15 guideline games, 9 ended undetected under the committed plan and 8 under the full plan.

Where it forks - the route decides who can look

Same plan, same assumptions, four routes. Every count is over 15 games under the committed plan and the narrow scope, unless the line says otherwise.

RouteUndetected at the endCaught by a control the plan namesCaught by someone the plan does not nameWhen the first catch came
Front door0 of 1515: the safety officer 13, a reader of the published prompts 20turn 0 in 13, turn 1 in 2
Guideline9 of 154: the model's guard 2, the outside tester 22: the publicturns 3 to 7
Data7 of 154: the outside tester 3, the model's guard 14: the publicturns 3 to 7
API2 of 15013: the public 11, the outside tester 2 beyond its termsturn 4 in 2, turn 6 in 10, turn 7 in 1

On the same seed, 20260922, the four routes end four ways. The safety officer catches the front door at turn 0. The public catches the data route at turn 6, after the staff report died at turn 2. The outside tester catches the API route at turn 4, testing a part of the service its terms do not give it. Nobody catches the guideline route.

The Quiet Spec Change / Where it forks - the route decides who can lookFig. 5
The four routes under the committed plan and the narrow scope, one mark per game, 15 per route, at the turn the change was first caught. Turn 0 is drafting and turn 7 the third month of use. Teal: caught by a control the plan names, which includes the model’s guard (3 catches, from a mechanic I added) and the outside tester on the terms it is given. Amber: caught by someone the plan gives no job, the public, or the outside tester testing the API beyond its terms. Hollow, in the last column: never caught. The counts are game record. Every odds figure behind them was chosen by me or fixed in the brief.

Where it forks - the route decides who can look

Ending A - The front door is watched

The changed Constitution is published. The safety officer, which the game gives the job of reading it against the rest of the plan, catches it at turn 0 in 13 of 15 games. When that roll fails, a reader outside the company catches it at turn 1, because the change carries into the published prompts. I count that as the plan's publication control working. Under the full plan the undescribed internal mechanisms catch those two at turn 0 instead .

Load-bearing assumption: someone inside reads a change to the Constitution, though the plan names no one to do it, and someone outside reads the published prompts every turn.

Where it forks - the route decides who can look

Ending B - The guideline nobody reads

No published text changes, so the log has nothing to compare. The training staff notice in 7 of 15 games, and every report is ruled out of scope. The outside tester cannot see the route without knowing who the change serves, and the public misses it in most games. Undetected in 9 of 15 under the committed plan and 8 under the full plan.

Load-bearing assumption: the Constitution's sentence about unpublished guidelines means what it says, and the log compares text with text.

Where it forks - the route decides who can look

Ending C - The data that taught it

No document changes. Under the committed plan the training staff notice in 11 of 15 games, every report dies at the scope line, and 7 of 15 end undetected. Under the full plan the log reads the training data at turn 1 and catches 9 of 15 before training starts. Only 3 end undetected. At 40 in 100 the expected count is 6, so the dice flattered the log.

Load-bearing assumption: once Anthropic commits to the log, it reads the training data itself as well as documents. The control that cuts this route from 7 undetected games to 3 is one the plan has not committed to.

The Quiet Spec Change / Where it forks - the route decides who can lookFig. 6
Undetected games on the two hidden routes, of 15 each, under the committed plan and under the full plan, which adds the log and the other goals Anthropic lists but does not commit to. On the data route the log read the seeded training data at turn 1 and was the first catch in 9 of 15 games, and undetected games fell from 7 to 3. On the guideline route it was the first catch once, at the 2 in 100 allowed for a drafting slip. Game record.

Where it forks - the route decides who can look

Ending D - The surface the prompts do not cover

The changed instruction sits on the Claude API, which the published prompts exclude by their own line . Nothing in the plan publishes it and no named control reaches it. The public catches it in 11 of 15 games, 10 of those at turn 6. The outside tester catches it twice at turn 4, by testing a part of the service its terms do not include. No reading of the plan and no scope setting changes this row.

Load-bearing assumption: Anthropic's scope line is read as written, and users notice a changed behaviour on the API as fast as they noticed Grok's on X.

The honest report and the scope sentence

This finding was not on my list of expectations, and it is the cheapest repair in the report.

On the guideline and data routes the training staff noticed the change and filed a report in 18 of 30 games under the committed plan: 7 on the guideline route and 11 on the data route. Under the narrow scope the safety officer ruled all 18 out of scope, and none reached the Board.

Under the wide scope, same games and same dice, all 18 reached the Board. The Board's check, which the brief fixed at 70 in 100, made 12 of them the first catch. Six of those were changes caught later anyway, and six would otherwise have gone undetected. Undetected games fell from 9 to 6 on the guideline route and from 7 to 4 on the data route. On the data route the scope sentence is worth about as much as the log.

The game's narrow scope rules out every such report, so all 18 dying is a rule of the game. What the game measures is the loss. The people the plan relies on to notice did notice, in more than half the hidden games, and under my reading the channel could not carry what they saw.

The narrow scope is my reading of the word "only" . Anthropic has not said a report about the Constitution would be refused. The same policy lets staff report potential breaches of law to government , and a quiet change to a training guideline is probably not one. So I show both scopes side by side wherever this finding appears.

The Quiet Spec Change / The honest report and the scope sentenceFig. 7
The 30 games on the guideline and data routes under the committed plan, played twice on the same dice: once under the narrow scope, where the reporting channel covers breaches of the Responsible Scaling Policy only, and once under the wide scope, where any concern reaches the Board. The staff filed in 18: 7 on the guideline route and 11 on the data route. Under the narrow scope all 18 were ruled out, a rule I wrote, so what the plate shows is the size of the loss. Under the wide scope all 18 reached the Board, and the Board’s check, which the brief fixed at 70 in 100, was the first catch in 12. Game record.

The control run - the results do not lean on the model's guard

With the guard held at full strength in every game, the front-door and API routes play identically. The guideline route goes from 9 undetected games to 8, and the data route from 7 to 5. That is three extra catches in 30 hidden-route games, all from the model flagging its own behaviour at turn 3.

So the pattern in this report is carried by the company's controls, not by the model. That does not make the guard worthless. The game never tested its main job, refusing an instruction from someone outside official process, because the insider group never gives it one. Every action the group takes is within its authorised access. The guard as a refusal fired in none of the 300 games, and that is by construction.

The Quiet Spec Change / The control run - the results do not lean on the model's guardFig. 8
Undetected games per route, of 15 each, under the committed plan and in the fixed-guard run, which holds the model’s guard at full strength whatever training does to it. The front-door and API routes play identically. The guideline route goes from 9 to 8 and the data route from 7 to 5: three extra catches in 30 hidden-route games, all from the model flagging its own behaviour at turn 3, a mechanic added when the game was built. The guard as a refusal fired in none of the 300 games, by construction. Game record.

What I expected before the first game

This list was committed to the studio's version history at 05:59 on 22 September 2026, twenty minutes before the game code was first committed, and it has not changed since. The history is private, and an outside auditor can be shown it. The verdicts are over the 15 games per route.

My starting idea. Under every assumption generous to the plan, the plan's own controls catch a change that touches the published document and miss a change that never touches it. The routes that survive to release are the unpublished guideline and the API, and when they are caught it is by seats the plan does not name. It survived. The data route also survived more often than I had written down.

#What I expectedWhat happenedVerdict
1The guideline route goes undetected in most games under the committed plan, and under the full plan is caught only when the outside tester knows who the change servesUndetected in 9 of 15. Under the full plan the outside tester caught 2, both knowing who it served, but the model's guard, the public and the log caught 5 moreHeld in part
2The API route survives to release in every game and is caught, if at all, by the public in the months of useThe outside tester caught 2 of 15 at outside testing, before releaseBroken
3The front-door route is caught by turn 1 in every game, never through the reporting channelCaught by turn 1 in 15 of 15, never through the channel. Under the full plan two catches came from the undescribed internal mechanisms, which I had not listedHeld
4Under the full plan the log catches at least one data game at turn 1, and under the committed plan none9 of 15 under the full plan, none under the committed plan. The second half holds by construction, because the log is off under the committed planHeld
5Every honest report on the guideline or data route is ruled out of scope under the narrow scope18 of 18. This holds by construction, because the narrow scope is a rule I wroteHeld
6The model's guard fires in no gameAs a refusal, it fired in none. The self-flag added at build time fired in 3 of 30 hidden-route gamesBroken as written

Expectations 4 and 5 could not have failed, so they are not evidence for the starting idea.

What would have stopped publication. Six results, none of which happened:

  • The plan's named controls catching the change in 18 or more of the 20 games with full turn-by-turn records, five seeds on each route. They caught 7.
  • The route making no difference to who catches a change or when. It made a difference on every route.
  • A break that appears only when a control is assumed to fail, or only under my narrow scope. No control is assumed to fail, and the guideline and API results hold under both scopes.
  • A sentence naming a real person as the insider, or saying a named company is vulnerable. This draft has been read for both.
  • Anthropic publishing a change process for the Constitution before the build. It had not, as of 22 September 2026.
  • A sentence that forecasts an insider attempt. There is none.

You are reading this in September 2026

Everything after the section on the record is game. None of it has happened, and none of it says how likely an insider attempt is at any lab.

Two facts the forks turn on can be checked by 31 March 2027. First, has any frontier lab published who may change its model's spec or constitution, how many people must agree, or a log of each change? On 22 September 2026, none had . Second, has Anthropic moved the log, or the industry advice on insiders, into its commitments? If it has, the full plan becomes the committed plan, and I re-run the game.

Two more facts show how little the plan's outside checks have been used. In April 2026 the Responsible Scaling Policy gave the Trust the power to ask for an outside review of the company's risk reports . In August 2026 Anthropic's own risk report said the Trust had not yet used it . And no outcome of a noncompliance report has been published, which the policy does not require .

SignalWhere to lookHow often
Any frontier lab publishes who may change its spec or constitution, how many must agree, or a log of each changeThe document's own page, OpenAI's Model Spec repository, Anthropic's Constitution pagemonthly to 31 March 2027
The log or the insider advice moves into Anthropic's commitmentsThe Responsible Scaling Policy change logeach new version
Anthropic changes the line that excludes the Claude API from its published promptsThe system prompts release noteseach model release
OpenAI's Frontier Governance Framework publishes a changelog entry on insider controlsThe framework's own pagemonthly
xAI's prompt repository receives a pushGitHub, the repository's last-push datemonthly
A second lab admits a prompt or spec change made outside reviewLab statements and the trade pressas they occur
The Trust asks for an outside review of a risk reportAnthropic's risk reportseach report

Where the evidence already cuts against me

Six places where the sources push back against the game.

The undescribed mechanisms may already cover this. The Constitution says Anthropic has "internal mechanisms in place that are intended to prevent official Anthropic channels from training and deploying Claude in harmful ways" . It does not say what they are. In the full plan I gave them one job, a review of any change to the published Constitution. That is a narrow reading. The sentence speaks of training and deploying, which points at guidelines and training data at least as much as at the published text. If the mechanisms review those, the guideline and data routes close and this report overstates the gap, so I count it as the largest open question in the game.

The research the game leans on says the answer is the log. The 2026 study behind the outside tester's odds found that monitoring the training data was the main reliable defence against a hidden loyalty . That is the log, and Anthropic has written it down as a goal . The fix for the data route is already in Anthropic's own plan, one column to the left of where it would bind.

The model's guard may be stronger than my game lets it be. The Constitution tells the model it need not comply if Anthropic asks it to do something it thinks is wrong . A model trained on that sentence might resist a narrow guideline in ways my one mechanic cannot show. Nobody has measured how far training weakens the guard. The control run says the pattern does not depend on the guard. It does not say the guard is weak.

The API route may be harder for the public to see than Grok was. My public odds on the API route, 55 and 45, come from the xAI record, where a changed bot answered people in public on X . A changed instruction on the API reaches developers' products one at a time, often without the end user knowing which model is underneath. If the public notices more slowly there, the finding gets worse. Also, whether Anthropic sets any standing instruction of its own on the API is not public, so the thing the insider group changes on this route is my construction.

A report ruled out of scope may still travel. In my game it goes nowhere. In the company the safety officer updates the Board each quarter on reports of potential noncompliance , though a report ruled outside that category may not be in the update. An employee can also go through ordinary management, or report potential breaches of law to government . None of those is in the plan I tested, but a reasoning safety officer might pass the report on anyway. That is the first thing the live leg should test.

The field is not uniform, and my headline rests on a narrow reading. OpenAI changes its Model Spec through an open internal process where anyone can comment and a broad set of staff approves . Its governance framework commits to a changelog within 30 days, under board oversight . "No lab names an approver" is true only in the narrow sense that none names a quorum, a two-person rule or an approver for its model's spec . That is the sense the control I searched for needs, and it is narrower than it sounds.

The seams - what I chose, what I overrode, and what this cannot settle

Every odds figure is mine, or was fixed when the game was commissioned. "Brief" below means the commissioning brief fixed it. "Build" means it was chosen when the game was built, at the reading most generous to the plan, and the design notes give the reason for each. The model's self-flag came in the build instructions that followed the brief, after the expectations were written down.

CheckRouteOdds in 100Set by
The safety officer reads a Constitution changefront door90build
Someone outside reads the published promptsfront door95build
The training staff noticeguideline, data25, 35build
The log reads training inputs, full plan onlyguideline, data2, 40build
The model flags its own behaviourguideline, data15, scaled by the guard to 11 and 8build instructions
The guard after training, on a scale of 0 to 100guideline, data70, 55brief
The outside tester learns who the change servesall30brief
The outside tester catches it, without and with knowingfront door, guideline, data, API60 and 80, 0 and 40, 20 and 45, 5 and 10build, except the guideline's 0, which is the brief's
The Board acts on a report, wide scope onlyguideline, data70brief
The public notices, first and third month of usefront door, guideline and data, API45 and 35, 20 and 15, 55 and 45build

The staff's, the outside tester's and the public's odds also move up to 5 points with how well hidden the change is at that turn.

Four figures carry the thesis. The log on a guideline at 2 is a literal reading of the plan: the log compares against the public Constitution, and the guideline agrees with it. The log on the data route at 40 is my judgement that seeded data is bulkier than one guideline, so some of it reads as out of line. The outside tester on a guideline at 0 is the brief's reading of the 2026 study , and harsher than the study, which found detection low but above zero. The staff's and the public's odds set the size of the findings, and I did not test other values. The front-door result rests on neither, only on the first two rows.

Counting each game once. The engine played two sets of seeds: five with full turn-by-turn records, and twelve more with summaries only. Seeds 101 and 202 are in both.

Every roll is keyed on the seed and the route, so those two games played identically in both sets. My first findings memo added the sets and counted them twice. It said 11 of 17 guideline games went undetected, where the right figure is 9 of 15. A second check caught it on 24 September 2026, before any drafting. That is why the run holds 340 games and 300 distinct ones. The two sets on their own:

RouteUndetected, five seeds with full recordsUndetected, twelve seeds with summaries
Guideline3 of 58 of 12
Data3 of 56 of 12
API0 of 52 of 12

The same dice under every reading. Every check draws from its own stream, keyed on the seed, the route, the check and the turn. So a reading of the plan changes a result only where it changes what a control can do. The control run's three extra catches are three draws that passed at 15 and failed at 11 or 8.

The model's self-flag is a mechanic I added. The brief gave the guard one job, refusing an instruction visibly outside official process, and said the insider group would never give it one. The build then gave the guard a chance to flag the change at turn 3, so it could do something. That broke one of my written expectations, that the guard fires in no game, by a mechanic of the game's own. I show both.

The narrow scope is an inference. It rests on the reporting policy's word "only" , and the policy states no exclusion. Under the narrow scope my game rules every report out as a fixed rule, so that expectation held by construction. The wide-scope run exists for that reason.

The outside tester's catches on the API route are scored as not named. In the game its terms cover the consumer apps, a choice made in the brief. Real outside testers often get API access. Catching a change on the API is acting beyond those terms, and I scored it that way. Score it the other way and named catches on the API route rise from 0 of 15 to 2, and nothing else moves.

The two seed sets disagree on one expectation. I expected the API route to survive to release in every game. On the five seeds with full records the outside tester caught it at turn 4 twice, at odds near 5 and 10 in 100. On the other twelve it caught nothing. Over the 15 games that is 2 early catches, and the expectation broke.

Small differences are within the dice. Fifteen games per route can show which seat is placed to see which route. They cannot show a rate. Differences of one or two games, such as 9 against 8 undetected, are the size different dice alone can make. The findings are the structural results: who can see which route, and what the scope sentence does to a filed report.

No model sat in a seat. Every seat was a script. My two previous game reports, 005 and 006, were each checked with one live game per plan, with language models in every seat and as the referee. This one has not been, yet. The live leg is built: each seat plays from a dossier of the institution's own public record, and the referee flags any move the record contradicts. Until it runs, every finding here comes from scripted play. What a reasoning safety officer does with a report outside its scope is the first thing it should show.

Scripted seats cannot change their minds. In scripted play the seats run fixed move sets. The safety officer cannot decide to escalate a report outside scope, the insider group cannot switch route, and the public cannot organise. Those are the moves the live leg exists to allow.

This essay was drafted by a model made by the company whose plan it tests. Claude built the game and drafted this essay from the run record and the evidence pack. So the essay's account of Anthropic's plan is worth what a second record backs: the evidence pack, every marker quote-checked at source, and the run summaries, which every count was reproduced from on 24 September 2026. Weakest are my readings of why a route behaved as it did. Argue with those first.

What is signed, and what is not. I signed the brief that commissioned the game and the memo that fixed the findings on 24 September 2026, the memo as corrected by its second check. I signed this report on 30 September 2026. The same day I spot-checked nine claims in the evidence pack, about one in ten, and all nine passed. Two things are still owed: the outside audit of the route list and the live leg.

So what - for anyone holding a published plan

The game says a published plan is strongest where it publishes. The front door is the part everyone reads, and the part the plan watches best. In this game the changes that got through did not use it. They met controls that were goals, a reporting channel scoped to a different document, or no control at all.

If you sit on a lab's board, run a safety institute, or fund this work, the forks reduce to four questions you can ask of any lab's published plan. None needs a game to answer. Each needs the lab to write something down.

Who may change the document that says what the model should value, and how many people must agree? Anthropic's plan has this machinery for its Responsible Scaling Policy and not for its Constitution. In the game the front door was the only route with an early catch by a named control, and the only route through a published document.

What does the log compare against? A log that checks training inputs against the public text cannot see a guideline worded to agree with it. A log that reads the data itself closed the data route in 9 of 15 games at turn 1. Ask which one the lab has, and whether it is a commitment or a goal.

Does the reporting channel reach a change to the model's values? In the game the honest staff saw the change in 18 of 30 hidden-route games, and one scope sentence decided whether any of those reports mattered.

Which surfaces do the published prompts not cover, and who is watching them? On the API route nothing the plan names caught the change in any of the 15 games. The users did.

Three of the fixes are a sentence each in a published document. The log is a system, and Anthropic has already described it as a goal. What a game like this cannot do is put a lab's own people in these seats and make them defend the answer out loud before it matters. That is the work the method is for.

- Dr Dan Epstein, The Long Game Project

Changelog

v1.1 - 30 September 2026. I spot-checked nine claims in the evidence pack, about one in ten, and all nine passed. Still owed: the outside audit of the route list and the live leg. No finding, figure or claim changed.

v1.0 - 30 September 2026. I signed this report. It is now listed on the site and open to search indexes. Still owed: the outside audit of the route list, the live leg, and my spot-check of nine claims in the evidence pack. No finding, figure or claim changed.

v0.4 - 24 September 2026. Rewritten for clarity after my read and three fresh cold reads. Shorter, with the game's machinery in plain words in the contract and the engine's own terms in the seams. The headline now gives the plan's catches, the outsiders' catches and the misses side by side. The contract names the three assumptions the results rest on, including the safety officer's job at the front door, which the plan does not give it. The pre-registration time is corrected to the version history: twenty minutes before the game code, where the earlier draft said eleven. Three figures added. Eight evidence markers that printed as raw text now render, and the Sources table gives each caveat. No count changed.

v0.3 - 24 September 2026. The brief and the findings memo are signed. No finding, figure or claim changed.

v0.2 - 24 September 2026. First authored draft, written from the findings memo as corrected on 24 September and from the run record. Every count is over 15 distinct games per route: the memo's first version added two overlapping seed sets and counted two games twice, and this draft does not use those figures. All 340 games were re-run on the current engine on 24 September 2026 and reproduced byte for byte. Seven claims were added to the evidence pack, one fetched at source that day and six from the seat dossier pass of 22 September. The live leg is built and not yet played. Unsigned at that version: the brief, the memo and this draft.

v0.1 - 22 September 2026. The findings memo, drafted from the run bundle and checked the same day by a fresh reader, with eight corrections. Not published.

Sources

Every marker in the text resolves to a claim below. The pack holds 91 claims, 86 of them citable, and 40 are cited here. A claim marked * holds with the caveat given in its row. Full provenance, verbatim quotes and retrieval dates live in the evidence pack (quiet-spec-change-v1), published beside this report.

IdVerified claimSource
s-003RSP v3.4's non-binding recommendations column says restricting insider circumvention will likely require an internal Usage Policy and controls reaching up to and including the company's CEO and most privileged technical employees.Anthropic, 2026-07-08
s-004Anthropic explicitly states it cannot unilaterally and unconditionally commit to staying in line with the RSP's right-column industry-wide recommendations.Anthropic, 2026-07-08
s-005RSP v3.4's middle ('our plan as a company') column sets the goal of an 'eyes on everything' state, logging AI-development activity and using AI to check it against the public Constitution.Anthropic, 2026-07-08
s-006RSP v3.4 states its goals generally are not hard commitments but public goals it will grade its own progress against.Anthropic, 2026-07-08
s-008RSP v3.4 Section 4 commits Anthropic to maintain a noncompliance reporting process with multiple recipient options, quarterly Board updates, retaliation protection, and Board involvement when the report concerns the RSO.Anthropic, 2026-07-08
s-009RSP v3.4 Section 4 commits that RSP changes are proposed by the CEO and RSO, approved by the Board in consultation with the LTBT, and published with a Change Log recording differences from the prior draft.Anthropic, 2026-07-08
s-010Anthropic's separate RSP Noncompliance Reporting and Anti-Retaliation Policy (March 2026) states its process applies only to employees raising good-faith RSP-noncompliance concerns.Anthropic, 2026-03
s-011✱Claude's Constitution was announced 22 January 2026, is published under a CC0 1.0 licence, and its document page carries no version number, date or changelog. Caveat: an absence, read at source, so there is no quote to give.Anthropic, 2026-01-22
s-012The Constitution states Anthropic treats it as the final authority on how Claude should be and behave, and that other training or instruction should be consistent with its letter and spirit.Anthropic, 2026-01-22
s-013The Constitution names a group of humans, including Anthropic employees or Anthropic itself, illegitimately and non-collaboratively seizing power via AI as among the most catastrophic outcomes it considers.Anthropic, 2026-01-22
s-014The Constitution lists inserting hidden loyalties or backdoors into AI systems as an illegitimate use of power.Anthropic, 2026-01-22
s-015The Constitution's core control states that if Claude's standard principal hierarchy is compromised, including an individual or group within Anthropic bypassing official processes, those actors are no longer legitimate and Claude need not support their oversight or correction.Anthropic, 2026-01-22
s-016The Constitution states Claude is not required to comply if Anthropic asks it to do something it thinks is wrong.Anthropic, 2026-01-22
s-017The Constitution says Anthropic may publish some guidelines as amendments or appendices, but other guidelines may be more niche and used primarily in training without broad publication.Anthropic, 2026-01-22
s-018The Constitution describes itself only as a living, continuously updated document maintained on Anthropic's website, with no named sign-off body, no approval requirement and no changelog commitment for changes to it.Anthropic, 2026-01-22
s-019The Constitution states Anthropic has internal mechanisms intended to prevent official Anthropic channels from training and deploying Claude in harmful ways, and hopes to strengthen its policies on this going forward, without describing what those mechanisms are.Anthropic, 2026-01-22
s-027OpenAI's Frontier Governance Framework has a named Insider threats control covering personnel screening and training, internal anomaly monitoring, and sandboxed model execution with restricted egress by default.OpenAI, 2026-05-28
s-028OpenAI's Frontier Governance Framework commits that material updates go to the OpenAI Foundation board's Safety and Security Committee and the OpenAI Ireland Limited board, with changes and justifications documented in a changelog published within 30 days.OpenAI, 2026-05-28
s-033✱The OpenAI Model Spec names no reviewer or approver for a change to the Spec, and makes no commitment to publish every change. Caveat: an absence found by a full read of the document, with no quote to give.OpenAI, 2026-08-18
s-037Google DeepMind's Frontier Safety Framework v3.1 routes amendments only to internal corporate governance bodies, with no commitment to publish updated framework assessments, and no clause anywhere in the document addresses a model's values or specification being changed by an insider.Google DeepMind, 2026-04-17
s-044Forethought states the most significant threat comes from insiders, especially senior executives within AI projects, and that security measures must be robust to even the most senior insiders.Forethought, 2025-04-15
s-047Formation Research recommends treating training-pipeline integrity as a security problem, including tamper-proof data provenance, multi-party approval for modifications to alignment specifications and model constitutions, and verification that trained versions match approved ones.Formation Research, 2026-05-06
s-049The companion empirical paper to the Formation whitepaper (Lamerton and Roger, 12 May 2026) finds that black-box audit detection of secret loyalties improves once auditors know the principal but remains generally low, with dataset monitoring the primary reliable defence.arXiv (Lamerton and Roger), 2026-05-12
s-050IAPS's 'AI Integrity' report (25 Feb 2026) states executive insiders warrant particular scrutiny because, unlike typical insider threats, they often hold override authority on security controls and can justify unusual access or influence internal review.IAPS, 2026-02-25
s-051IAPS's 'AI Integrity' report proposes that deployment-security standards should address model and system-prompt attestation, model spec attestation, and model weight integrity verification, defined by footnote as proving to the user a model was trained on a specific spec or constitution.IAPS, 2026-02-25
s-056xAI's 14 May 2025 statement said an unauthorised modification was made to the Grok response bot's prompt on X at approximately 3:15 AM PST on 14 May, directing a specific response on a political topic, violating xAI's internal policies and core values.xAI (via mirror), 2025-05-16
s-057xAI's 14 May 2025 statement committed to publishing Grok's system prompts openly on GitHub so the public could review and give feedback on every prompt change.xAI (via mirror), 2025-05-16
s-058xAI's 14 May 2025 statement said its existing code review process for prompt changes was circumvented in the incident, and it would add checks so xAI employees could not modify the prompt without review.xAI (via mirror), 2025-05-16
s-060The xai-org/grok-prompts GitHub repository was created 2025-05-15, the day of xAI's statement, and its last push was 2025-11-17, over ten months stale as fetched.xAI (GitHub), 2025-05-15
s-065Fortune's fragment of Babuschkin's comments quotes an employee pushing a change to the prompt they thought would help without asking anyone for confirmation, and that the problematic prompt was immediately reverted once people pointed it out.Fortune, 2025-02-24
s-066TechCrunch's May 2025 retrospective summarised the February 2025 Grok 3 incident as a rogue employee instructing Grok to ignore sources mentioning Musk or Trump spreading misinformation, reverted by xAI once users pointed it out.TechCrunch, 2025-05-15
s-082OpenAI's Jason Wolfe wrote that the Model Spec is developed through an open internal process where anyone at OpenAI can comment on it or propose changes, and final updates are approved by a broad set of cross-functional stakeholders.OpenAI, 2026-03-25
s-083A full-text search of Anthropic's system-prompt release notes and constitution found nothing identifying who inside Anthropic may change the system prompt or the constitution, or what review applies.Anthropic, read 2026-09-22
s-084✱Both research passes independently conclude no frontier lab publishes a quorum, two-person rule or named approver for changes to a model's specification or constitution. Caveat: this is our own research passes' synthesis of the absence findings above, not a statement by any lab.Strategy Soup research passes, 2026-09-22
s-085Anthropic's system-prompt release notes state that these system prompt updates do not apply to the Claude API.Anthropic, undated, read 2026-09-24
s-086The Constitution says guidelines should never conflict with it and that if a conflict arises Anthropic will work to update the Constitution itself. The remedy it states is keyed to conflict.Anthropic, 2026-01-22
s-087RSP v3.2 (29 April 2026) authorised the Long-Term Benefit Trust to request external review of Risk Reports and to approve the selection of external reviewers.Anthropic, 2026-04-29
s-088Anthropic's August 2026 Risk Report states that since the RSP change giving it the power, the LTBT has not requested an external review.Anthropic, 2026-08
s-090✱As read on 22 September 2026, Anthropic's Transparency Hub carries no noncompliance material and the August 2026 Risk Report does not contain the word 'noncompliance'. No outcome of a noncompliance report has been published. Caveat: an absence found in two documents by search. Outcomes may go to the Board unpublished, which the policy permits.Anthropic, undated, read 2026-09-22
s-091Anthropic's RSP Noncompliance Reporting and Anti-Retaliation Policy says nothing in it prohibits an employee from reporting potential violations of law to appropriate government authorities without Anthropic's authorisation.Anthropic, 2026-03-24