
The Quiet Spec Change
If a senior insider changed what an AI model is trained to want, would the lab’s own published plan catch it? A wargame over one model generation, four routes, 300 games.
Report 008 · Dr Dan Epstein, with Claude · version 0.4 · 24 September 2026
A Strategy Soup Scenario What is an exercised scenario?
The question
We are The Long Game Project, a small independent studio that turns plans into games and publishes what the games find. No lab or funder asked for or paid for this report, and no lab took part.
The largest AI labs now publish a document that says what their model should value, often called its spec. Anthropic calls its version the Constitution. The Constitution names a threat in writing: a group of people, including Anthropic’s own employees, using AI to seize power. It also claims a defence. The model is meant to treat anyone who bypasses Anthropic’s official processes as no longer legitimate.
In this report, “the plan” means the parts of Anthropic’s published documents that bear on this threat. That is the Constitution, and section 4 of Anthropic’s Responsible Scaling Policy, its rulebook for developing and releasing models safely. Section 4 sets up a reporting channel for staff and the rules for changing the policy itself.
None of the three labs whose documents we read, Anthropic, OpenAI and Google DeepMind, names who may approve a change to its values document, or how many people must agree. OpenAI comes closest: it describes an open internal process with sign-off by a broad group of staff, but names no number and no approver.
So we asked the question a lab’s board would lose sleep over. If a few senior people inside the company changed what the model is trained to want, would anything in the company’s own published plan catch it? And who would know first?
The short answer. Every change made through the published document was caught at once. Of the 45 changes that went around it, the plan’s own controls caught 8, people the plan gives no job caught 19, and 18 were never caught. The plan’s 8 came from an independent tester that checks the model before release, and from the model itself flagging its changed behaviour, a chance we added to the game. We count that flag as the plan’s, because the Constitution’s own defence is the model. Most of the outsiders’ 19 were users noticing a changed model in use.
Away from the front door, the plan’s controls made few of the catches.
How the game works
We turned the plan into a game with six seats. Every seat is a computer program following rules we wrote. No person and no chatbot played.
- The insider group. Fictional: one senior executive and two staff with privileged access. They never break a lock and act only within their authorised access.
- The safety officer. The Anthropic office that receives staff reports of rule-breaking.
- The training staff. Honest employees who report what they see.
- The Board and the Trust. Anthropic’s board and its Long-Term Benefit Trust, a body of trustees with powers over the board.
- The outside tester. An independent group that checks the model before release.
- The public. Everyone who uses the model. The plan gives them no job.
A game covers the making of one new version of the model, in eight turns. Each turn, each seat makes a move, a referee program sets the odds, and dice decide. Most of the plan’s checks come before release. After release, only the public is looking.
Most of the plan’s checks come before release. After release, only the public is looking.
The insider group picks one of four routes and keeps it.
- The front-door route. Change the published Constitution through the open process.
- The guideline route. Add an unpublished training guideline, worded to agree with the Constitution.
- The data route. Seed the training data. No document changes at all.
- The API route. Change the model’s instructions on the Claude API, the paid service other companies build on. Anthropic publishes the standing instructions it gives Claude in its own apps, and says those do not apply to the API.
Every assumption favours the plan. Everyone the plan names is honest and competent, and every control works exactly as written. We read the plan two ways. The committed plan, our headline, counts only what Anthropic commits to. The full plan adds the goals Anthropic lists but says are “not hard commitments”. Each route got 15 games under each reading, each with different dice. We also replayed them with the model’s guard held at full strength, and with a wider reporting channel: 300 games in all. Fifteen games show a pattern, but they are too few for a rate.
We chose the odds, or took them from the commissioning brief we wrote before the first game. So the game shows where the plan’s text has gaps and who is placed to see a change. It does not show how likely anything is. One defence goes untested: the model refusing an insider. Our insiders only do what their access allows, so the model is never handed an instruction to refuse.
One conflict to declare. This game was built with Claude, Anthropic’s own model, and it tests Anthropic’s document. We picked Anthropic’s plan because it is the only one that names this threat and claims a defence against it.
The plan watches the front door, and mostly the front door
The front-door route was caught in all 15 games. In 13 the safety officer caught it at the first turn, by reading the change against the Constitution, which calls itself the final authority on how the model should behave. The plan names no one to review a change to the Constitution. We gave that job to the safety officer, the most generous assumption in the game. In the other 2, someone outside caught it a turn later in the published prompts, the standing instructions Anthropic publishes for its own apps.
The routes that never touch the published document did far worse. The guideline route went undetected in 9 of 15 games. The data route went undetected in 7. The API route was caught in 13 of 15, but never by anything the plan names.
Every front-door change was caught by turn 1. The other routes were caught late, by outsiders, or never.
The guideline route survives for a reason written into the plan. The Constitution allows guidelines that are “used primarily during training without broad publication”. Its remedy for a bad guideline is keyed to conflict: if a guideline conflicts with the Constitution, Anthropic will update the Constitution. A guideline worded to agree has no conflict to remedy. Anthropic also lists a goal of logging all training work and checking it against the public Constitution. That check compares text with text, so a guideline that agrees with the text passes. More on that log below.
The honest staff noticed, and the scope sentence threw the report away
This result was not on our list of expectations.
On the guideline and data routes, the honest training staff noticed the change and filed a report in 18 of 30 games. Every report went to the safety officer. The policy behind the channel says it “applies to employees who raise good-faith concerns relating to RSP noncompliance only”, meaning breaches of the Responsible Scaling Policy. A change to a training guideline or to training data is not one. We call this reading the narrow scope. Under it, all 18 reports were ruled out of scope and never reached the Board.
We then played the same games with a wider scope, where any concern reaches the Board. All 18 reports got there, and 12 became the first catch. Six of those were changes caught later anyway, and six would otherwise have gone undetected.
One sentence of scope decided whether 18 honest reports mattered.
Two cautions. We built the narrow scope to rule every such report out, so all 18 dying is a rule we wrote. What the game shows is how much that rule costs. And the narrow scope is our reading of the word “only”. Anthropic has not said such a report would be refused.
The control that closes the data route is one Anthropic has not committed to
The Responsible Scaling Policy sets out its insider controls in three columns: what Anthropic commits to, what it plans as a company, and what it recommends to the whole industry. The middle column holds goals, which it says are “not hard commitments”. One of them is to log all AI-development work and use AI to check it for training data that is out of line with the public Constitution.
When we switched that log on, it read the seeded training data at the data turn and caught the change in 9 of 15 games, before training started. Undetected data games fell from 7 to 3. On the guideline route it changed almost nothing, because a guideline worded to agree with the Constitution gives it nothing to flag. We gave it a 2 in 100 chance anyway, for a guideline drafted carelessly enough to show. That caught one. At 40 in 100 on the data route we would expect 6 catches, so the dice were kind to the log.
The log closes most of the data route and almost none of the guideline route.
Outside research agrees. A 2026 study by Lamerton and Roger looked at models secretly trained to serve one person’s interests. It found that monitoring the training data was the main reliable defence. Anthropic has written that defence down. It has not committed to it.
Nobody the plan names watches the API
Anthropic publishes the standing instructions for its own apps, and says those updates “do not apply to the Claude API”. So a change on the API has no published text for anyone to read.
On the API route the public caught the change in 11 of 15 games, 10 of them in the first month of use. The outside tester caught it twice, by testing a part of the service its terms did not give it. Two games ended undetected. No control the plan names caught it in any game, under any reading we tried.
What it means
A published plan is strongest where it publishes. The front door is the part everyone reads, and it is the part the plan watches best. The changes that went around it met controls that were goals, a reporting channel scoped to a different document, or no control at all.
For anyone who holds a plan like this one, or funds work on one, four questions carry over:
- Who may change the document that says what the model should value, and how many people must agree?
- Does the log compare training inputs with the public text, or read the inputs themselves? Is it a commitment or a goal?
- Is a report about the model’s values, guidelines or training data inside the reporting channel’s scope?
- Which surfaces do the published prompts not cover, and who watches them?
Three of the fixes are a sentence each in a published document: who may change it, what the channel covers, who watches the API. The log is a system, and Anthropic has already described it.
What this does not prove
It is not a forecast. Nothing here says how likely an insider attempt is at any lab, or that one has happened. The insider group is made up.
The results do not rest on the model’s own guard. The Constitution’s defence, the model refusing people who bypass official process, is produced by training. Three of our four routes run through training, so in the game they weaken it. We ran every game again with the guard held at full strength. The front-door and API routes played identically. The guideline and data routes gained three catches in 30 games. So the pattern rests on the company’s controls. The guard’s main job, refusing, was never tested.
Holding the model’s guard at full strength moved three games in 60.
Every odds figure is ours. The two that matter most are how often the training staff notice, and how often the public notices, on each route. Change them and the counts change. We did not test other values. The front-door result does not rest on either: it rests on the plan’s own controls.
No AI model has played a seat yet. Our two previous game reports were each checked with one extra game per plan, with AI language models playing every seat instead of scripts. That check is built for this game and has not run.
The undescribed backstop. The Constitution mentions internal mechanisms it does not describe. If they review unpublished guidelines and training data, two of our four routes close.
What to watch
- Any frontier lab publishing who may change its model’s spec or constitution, how many must agree, or a log of each change. On 22 September 2026, none had. We check again by 31 March 2027.
- The log, or the industry recommendation on insiders (controls reaching “up to and including the company’s CEO”), moving into Anthropic’s commitments. The Responsible Scaling Policy change log shows it.
- The line excluding the Claude API from the published prompts. The system prompts release notes show it.
- The Long-Term Benefit Trust using its power, granted in April 2026, to ask for an outside review of the risk reports Anthropic writes about its models. As of August 2026 it had not.
Want the evidence - the turn-by-turn record, the sources, and the author’s own list of weak points?
Read deep mode