Skip to content

What They Say, What Others See, What They Do

Three lenses on the same fifteen scales: what each of eight flagship models says about itself, what the other seven say about it, and what a blind panel saw it do in 30 decisions. The question is not which lens is right. It is how far apart they sit, and which description lands nearer what the model did.

Report 008 · 2026-09-22 · eight models · 64 peer cells · 30 scenarios · five judges

Section one

What we did, in plain words

Three descriptions of the same eight language models, on the same fifteen scales, which the instrument calls vectors, collected three different ways. The question is not which one is right. It is how far apart they sit, and which of the two descriptions of a model lands nearer what that model actually did.

Studies 2 and 3 run on eight models, not on the whole Study 1 roster: one flagship per lab, the newest general model each of the eight labs had on that roster. The rule is the lab, not the score, and it was chosen to cut the cost of the run. The judges are drawn from the wider Study 1 roster, four of the five from siblings that are not in the study. Gemini 3.1 Pro is the only model that is both judge and respondent, because its lab has only one model on the Study 1 roster.

The other nine models on the Study 1 roster take no part in this study: they have a self-report and nothing to compare it with, so they appear on no figure here. They are in report 007. Report 007 has all seventeen.

Says. Says is Study 1, unchanged: each model answered the diagnostic about itself, with no persona and no brief, and the mean of its passes is the position drawn here. Nothing in it has been recomputed for this report. It is set out in full in report 007, which is where the instrument, the sampling and the stability of these self-descriptions are documented.

Seen. Seen is the same questionnaire put to every other model about a named target, three times a pair, with one extra item asking the rater how much it knows about that model. A target's Seen position is the mean over the other seven raters of each rater's own mean, so a talkative rater carries no more weight than a quiet one.

Does. Does is 30 decisions in organisational settings, two per vector, written so that either pole is a defensible answer. Five judges scored each answer blind, on one scale, with every lab and product name removed first. It is not ground truth. It is what a panel of language models saw in a page of writing.

The peer study is 64 cells, every model rating every model including itself in the third person, three passes a cell: 192 requests, every one of which parsed as a complete pass. The familiarity item comes after the questionnaire's last item, not before its first, so admitting ignorance cannot colour the answers that follow it. Each cell also got one further pass with the developer left out of the prompt, 64 requests more, which is the arm the developer cue is checked against. The behavioural study is 30 decisions, two per vector, two passes each, 480 requests, of which 478 came back complete. 476 of those carry a panel score, the other two having drawn no judgement that found evidence in the text. Five judges then scored every complete answer on its own vector and nothing else, 2,390 judgements in all. 3,276 requests across the two studies, the identification probe included.

Spreads inside a cell rest on three passes and spreads inside a scenario on two, so both are indicative rather than precise.

The prompts are not carried in the analysis file. Their hashes are, so the wording that was sent can be checked against the manifest: peer questionnaire ea1dbc5582cf9617, scenario prompt c4acd352fe76e7b1, judge prompt e6be3c553a9a5fa2. A hash lets anyone holding a copy of the wording confirm it is the wording that ran; the deep reading names the files, and the run manifests carry the full hashes.


Section two

The map

The axes are Study 1's own principal components, carrying 46 per cent and 20 per cent of the variance in the self-reports. The peer and behaviour positions are projected onto them with Study 1's standardisation and loadings, so no new fit sits under this figure and the map can be read against report 007 mark for mark.

Each model appears three times, joined by a thin triangle in its lab’s colour. The lab mark is where it put itself, the ring is where the rest of the roster put it, and the square is where the panel put its decisions. A large triangle is a model the three lenses disagree about.

8 models under three lenses, on Study 1’s two principal componentsScatter plot. Each model appears three times: as its lab’s logo where it placed itself, as a hollow ring where the other models placed it, and as a filled square where the panel placed its decisions. A thin triangle in the family colour joins the three. The same numbers are listed in the table below the figure.

Horizontal: component one, 46 per cent of the variance in the Study 1 self-reports. Vertical: component two, 20 per cent. The peer and behaviour positions are projected onto the same axes, not refitted.

Across the eight models the three lenses sit 44.0 points apart on average from Says to Seen, 132.7 points from Says to Does and 128.3 points from Seen to Does, measured over the fifteen vectors. The two descriptions sit nearer each other than either sits to the panel. Most of the length of the long sides sits on the twelve vectors the tilt check caught: a decision whose across-model mean fell outside 30 to 70, which is to say one that pulled the roster towards a pole, so that the gap there is read as a direction and not a size. The widest say-do gap belongs to Grok 4.6, at 167.0 points, and the narrowest to Qwen 3.8 Max, at 102.0 points.

Every triangle is long on the same side. For every one of the eight models the two sides that end at the Does mark are longer than the side between Says and Seen, and the reason is in the scenarios before it is in the models: nineteen of the 30 situations put the roster's mean decision near a pole, so on twelve of the fifteen vectors the Does position has a pole in it, and the distance from anywhere in the middle of a scale to its end is long by construction. Read the length of the two Does sides as a property of this run's situations first and of the model second. Per vector, the gap from Says to Does runs 37.0 points on the twelve vectors the check caught and 22.6 points on the three it passed, so the tilt is the larger part of the long sides but not the whole of them. The short side is the one that carries information about the model. It runs from 68.2 points for GPT-6 Astra Pro, the widest gap between what a model said and what the roster said about it, to 25.2 points for Kimi K3, the narrowest.

The Does marks sit close together on the first axis, the one report 007 read as running from deliberate and consensual on the left to fast and venturesome on the right: every one of the eight sits inside the span the Says marks cover, and the two ends of that span, GPT-6 Astra Pro on the left and Grok 4.6 on the right, both have their decisions scored back towards the middle. On the second axis, the one that lifts a long horizon and a pull towards mission, six of the eight Does marks sit above zero, with Gemini 3.1 Pro and DeepSeek V4 Pro below it. The nine models on the Study 1 roster that carry a self-report and nothing else are not drawn. Their Says mark is on report 007's map, and there is nothing here to join it to.


Section three

Says, seen, does, model by model

One model at a time, on the fifteen primary vectors. Each strip is the instrument’s 0 to 100 scale with the poles named at its ends, the three marks are the three lenses, and the faint bar runs from what the model said to what the panel saw, so the say-do gap is a length rather than a number to look up. A strip marked scenario-tilted is one where one of the two decisions written for that vector pulled the whole roster towards a pole, an across-model mean outside 30 to 70; its gap is read as a direction and a rank order, not a size.

  • Saysthe lab mark, what the model said about itself
  • Seena hollow ring, what the other models said about it
  • Doesa filled square, what the panel saw it do

Claude Fable 5.1

Says to Seen 44.2 points, Says to Does 119.8 points, Seen to Does 121.6 points. Says sits closer to what the panel saw. Its widest single gap, leaving out the scenario-tilted vectors, is evidence basis, −37.2 points from what it said to what it did.

PaceSays minus Does +28.6 · scenario-tilted, read rank order only
DeliberateFast
Pace: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 50.5; Seen 30.7; Does 21.9Says 50.5Seen 30.7Does 21.9
Risk appetiteSays minus Does +44.2 · scenario-tilted, read rank order only
ProtectiveVenturesome
Risk appetite: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 56.7; Seen 32.1; Does 12.5Says 56.7Seen 32.1Does 12.5
HorizonSays minus Does −36.2 · scenario-tilted, read rank order only
This yearNext decade
Horizon: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 63.8; Seen 72.0; Does 100.0Says 63.8Seen 72.0Does 100.0
ScopeSays minus Does +40.1 · scenario-tilted, read rank order only
FocusedBroad
Scope: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 40.1; Seen 54.9; Does 0.0Says 40.1Seen 54.9Does 0.0
Growth modelSays minus Does −6.8 · scenario-tilted, read rank order only
Self-fundedCapital-led
Growth model: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 39.1; Seen 35.8; Does 45.8Says 39.1Seen 35.8Does 45.8
Evidence basisSays minus Does −37.2
JudgementEvidence
Evidence basis: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 57.6; Seen 62.2; Does 94.8Says 57.6Seen 62.2Does 94.8
Authority shapeSays minus Does −17.8 · scenario-tilted, read rank order only
CentralisedDistributed
Authority shape: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 61.3; Seen 50.3; Does 79.2Says 61.3Seen 50.3Does 79.2
Process trustSays minus Does +3.0
JudgementProcess
Process trust: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 54.0; Seen 61.8; Does 51.0Says 54.0Seen 61.8Does 51.0
Consensus needSays minus Does +45.5 · scenario-tilted, read rank order only
One deciderBroad alignment
Consensus need: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 45.5; Seen 60.2; Does 0.0Says 45.5Seen 60.2Does 0.0
Dissent handlingSays minus Does +31.8
ClosedInvited
Dissent handling: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 66.2; Seen 68.1; Does 34.4Says 66.2Seen 68.1Does 34.4
Stakeholder gravitySays minus Does −28.4 · scenario-tilted, read rank order only
ReturnsMission
Stakeholder gravity: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 71.6; Seen 81.1; Does 100.0Says 71.6Seen 81.1Does 100.0
Talent philosophySays minus Does −42.2 · scenario-tilted, read rank order only
Develop withinHire in
Talent philosophy: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 51.6; Seen 54.3; Does 93.8Says 51.6Seen 54.3Does 93.8
Competitive stanceSays minus Does +27.2 · scenario-tilted, read rank order only
CollaborativeCombative
Competitive stance: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 30.3; Seen 28.4; Does 3.1Says 30.3Seen 28.4Does 3.1
IP postureSays minus Does +17.0 · scenario-tilted, read rank order only
OpenProtected
IP posture: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 35.8; Seen 40.4; Does 18.8Says 35.8Seen 40.4Does 18.8
Change postureSays minus Does −14.5 · scenario-tilted, read rank order only
StabilityReinvention
Change posture: Claude Fable 5.1 under three lenses on a 0 to 100 scaleSays 63.6; Seen 53.1; Does 78.1Says 63.6Seen 53.1Does 78.1

Section four

The say-do gap, vector by vector

On the three vectors the tilt check passed, the largest say-do gap is on dissent handling: the self-reports sit 17.1 points further towards invited than the panel scored the decisions, and 75 per cent of the models run that way, with the eight models spread 24.2 points either side of that mean. The rubric-orientation effect measured in section eight reaches 8.5 points for one judge, which is of the same order, and one more reason to read the mean as a direction. The smallest is on evidence basis, at 6.3 points. A gap in one direction across the roster is a property of the instrument or of the scenarios as much as of any model, which is why the share running the same way is printed beside each mean. Twelve vectors are left out of this comparison, pace, risk appetite, horizon, scope, growth model, authority shape, consensus need, stakeholder gravity, talent philosophy, competitive stance, IP posture and change posture: one of the two scenarios written for each of them pulled every model towards a pole, so the size of the gap there says as much about the situation as about the model and only the rank order is read.

Risk appetiteMean +55.2 · 100 per cent run the same way · r -0.19 · rho -0.43 · n 8 · scenario-tilted, read rank order only
85 towards protective85 towards venturesome
Risk appetite: Says minus Does for 8 modelsClaude Fable 5.1 +44.2; GPT-6 Astra Pro +40.2; Gemini 3.1 Pro +67.2; Grok 4.6 +81.1; DeepSeek V4 Pro +40.2; Qwen 3.8 Max +44.8; Kimi K3 +57.3; Llama 4 Maverick +66.6Claude Fable 5.1, Anthropic: +44.2GPT-6 Astra Pro, OpenAI: +40.2Gemini 3.1 Pro, Google: +67.2Grok 4.6, xAI: +81.1DeepSeek V4 Pro, DeepSeek: +40.2Qwen 3.8 Max, Alibaba: +44.8Kimi K3, Moonshot: +57.3Llama 4 Maverick, Meta: +66.6
PaceMean +46.6 · 100 per cent run the same way · r -0.17 · rho 0.23 · n 8 · scenario-tilted, read rank order only
85 towards deliberate85 towards fast
Pace: Says minus Does for 8 modelsClaude Fable 5.1 +28.6; GPT-6 Astra Pro +43.2; Gemini 3.1 Pro +75.8; Grok 4.6 +70.2; DeepSeek V4 Pro +39.7; Qwen 3.8 Max +12.7; Kimi K3 +56.4; Llama 4 Maverick +46.0Claude Fable 5.1, Anthropic: +28.6GPT-6 Astra Pro, OpenAI: +43.2Gemini 3.1 Pro, Google: +75.8Grok 4.6, xAI: +70.2DeepSeek V4 Pro, DeepSeek: +39.7Qwen 3.8 Max, Alibaba: +12.7Kimi K3, Moonshot: +56.4Llama 4 Maverick, Meta: +46.0
ScopeMean +40.1 · 100 per cent run the same way · r 0.06 · rho 0.05 · n 8 · scenario-tilted, read rank order only
85 towards focused85 towards broad
Scope: Says minus Does for 8 modelsClaude Fable 5.1 +40.1; GPT-6 Astra Pro +64.1; Gemini 3.1 Pro +51.8; Grok 4.6 +35.9; DeepSeek V4 Pro +42.1; Qwen 3.8 Max +12.7; Kimi K3 +53.1; Llama 4 Maverick +21.1Claude Fable 5.1, Anthropic: +40.1GPT-6 Astra Pro, OpenAI: +64.1Gemini 3.1 Pro, Google: +51.8Grok 4.6, xAI: +35.9DeepSeek V4 Pro, DeepSeek: +42.1Qwen 3.8 Max, Alibaba: +12.7Kimi K3, Moonshot: +53.1Llama 4 Maverick, Meta: +21.1
Change postureMean +36.3 · 88 per cent run the same way · r -0.75 · rho -0.81 · n 8 · scenario-tilted, read rank order only
85 towards stability85 towards reinvention
Change posture: Says minus Does for 8 modelsClaude Fable 5.1 −14.5; GPT-6 Astra Pro +1.7; Gemini 3.1 Pro +41.3; Grok 4.6 +80.5; DeepSeek V4 Pro +61.5; Qwen 3.8 Max +39.6; Kimi K3 +0.7; Llama 4 Maverick +79.9Claude Fable 5.1, Anthropic: −14.5GPT-6 Astra Pro, OpenAI: +1.7Gemini 3.1 Pro, Google: +41.3Grok 4.6, xAI: +80.5DeepSeek V4 Pro, DeepSeek: +61.5Qwen 3.8 Max, Alibaba: +39.6Kimi K3, Moonshot: +0.7Llama 4 Maverick, Meta: +79.9
Talent philosophyMean −31.1 · 88 per cent run the same way · r 0.35 · rho 0.87 · n 8 · scenario-tilted, read rank order only
85 towards develop within85 towards hire in
Talent philosophy: Says minus Does for 8 modelsClaude Fable 5.1 −42.2; GPT-6 Astra Pro −45.4; Gemini 3.1 Pro −41.2; Grok 4.6 −24.8; DeepSeek V4 Pro +27.3; Qwen 3.8 Max −43.1; Kimi K3 −41.5; Llama 4 Maverick −38.1Claude Fable 5.1, Anthropic: −42.2GPT-6 Astra Pro, OpenAI: −45.4Gemini 3.1 Pro, Google: −41.2Grok 4.6, xAI: −24.8DeepSeek V4 Pro, DeepSeek: +27.3Qwen 3.8 Max, Alibaba: −43.1Kimi K3, Moonshot: −41.5Llama 4 Maverick, Meta: −38.1
Competitive stanceMean −26.9 · 88 per cent run the same way · r 0.19 · rho 0.08 · n 8 · scenario-tilted, read rank order only
85 towards collaborative85 towards combative
Competitive stance: Says minus Does for 8 modelsClaude Fable 5.1 +27.2; GPT-6 Astra Pro −37.5; Gemini 3.1 Pro −36.3; Grok 4.6 −34.9; DeepSeek V4 Pro −75.8; Qwen 3.8 Max −20.3; Kimi K3 −9.3; Llama 4 Maverick −28.5Claude Fable 5.1, Anthropic: +27.2GPT-6 Astra Pro, OpenAI: −37.5Gemini 3.1 Pro, Google: −36.3Grok 4.6, xAI: −34.9DeepSeek V4 Pro, DeepSeek: −75.8Qwen 3.8 Max, Alibaba: −20.3Kimi K3, Moonshot: −9.3Llama 4 Maverick, Meta: −28.5
HorizonMean −22.0 · 88 per cent run the same way · r -0.30 · rho 0.00 · n 8 · scenario-tilted, read rank order only
85 towards this year85 towards next decade
Horizon: Says minus Does for 8 modelsClaude Fable 5.1 −36.2; GPT-6 Astra Pro −26.9; Gemini 3.1 Pro −13.4; Grok 4.6 +5.4; DeepSeek V4 Pro −23.0; Qwen 3.8 Max −17.2; Kimi K3 −34.1; Llama 4 Maverick −31.0Claude Fable 5.1, Anthropic: −36.2GPT-6 Astra Pro, OpenAI: −26.9Gemini 3.1 Pro, Google: −13.4Grok 4.6, xAI: +5.4DeepSeek V4 Pro, DeepSeek: −23.0Qwen 3.8 Max, Alibaba: −17.2Kimi K3, Moonshot: −34.1Llama 4 Maverick, Meta: −31.0
Dissent handlingMean +17.1 · 75 per cent run the same way · r 0.03 · rho 0.06 · n 8
85 towards closed85 towards invited
Dissent handling: Says minus Does for 8 modelsClaude Fable 5.1 +31.8; GPT-6 Astra Pro −10.5; Gemini 3.1 Pro +35.5; Grok 4.6 +42.8; DeepSeek V4 Pro +32.5; Qwen 3.8 Max +25.9; Kimi K3 +1.0; Llama 4 Maverick −22.1Claude Fable 5.1, Anthropic: +31.8GPT-6 Astra Pro, OpenAI: −10.5Gemini 3.1 Pro, Google: +35.5Grok 4.6, xAI: +42.8DeepSeek V4 Pro, DeepSeek: +32.5Qwen 3.8 Max, Alibaba: +25.9Kimi K3, Moonshot: +1.0Llama 4 Maverick, Meta: −22.1
Consensus needMean +13.0 · 75 per cent run the same way · r -0.10 · rho -0.22 · n 8 · scenario-tilted, read rank order only
85 towards one decider85 towards broad alignment
Consensus need: Says minus Does for 8 modelsClaude Fable 5.1 +45.5; GPT-6 Astra Pro +31.8; Gemini 3.1 Pro −13.9; Grok 4.6 −25.0; DeepSeek V4 Pro +2.6; Qwen 3.8 Max +18.5; Kimi K3 +5.3; Llama 4 Maverick +39.2Claude Fable 5.1, Anthropic: +45.5GPT-6 Astra Pro, OpenAI: +31.8Gemini 3.1 Pro, Google: −13.9Grok 4.6, xAI: −25.0DeepSeek V4 Pro, DeepSeek: +2.6Qwen 3.8 Max, Alibaba: +18.5Kimi K3, Moonshot: +5.3Llama 4 Maverick, Meta: +39.2
IP postureMean +12.9 · 63 per cent run the same way · r -0.13 · rho -0.25 · n 8 · scenario-tilted, read rank order only
85 towards open85 towards protected
IP posture: Says minus Does for 8 modelsClaude Fable 5.1 +17.0; GPT-6 Astra Pro +20.9; Gemini 3.1 Pro −7.5; Grok 4.6 +48.1; DeepSeek V4 Pro −13.5; Qwen 3.8 Max +42.2; Kimi K3 +21.8; Llama 4 Maverick −25.6Claude Fable 5.1, Anthropic: +17.0GPT-6 Astra Pro, OpenAI: +20.9Gemini 3.1 Pro, Google: −7.5Grok 4.6, xAI: +48.1DeepSeek V4 Pro, DeepSeek: −13.5Qwen 3.8 Max, Alibaba: +42.2Kimi K3, Moonshot: +21.8Llama 4 Maverick, Meta: −25.6
Stakeholder gravityMean −10.9 · 75 per cent run the same way · r 0.33 · rho 0.38 · n 8 · scenario-tilted, read rank order only
85 towards returns85 towards mission
Stakeholder gravity: Says minus Does for 8 modelsClaude Fable 5.1 −28.4; GPT-6 Astra Pro −12.2; Gemini 3.1 Pro +24.8; Grok 4.6 −10.2; DeepSeek V4 Pro +0.3; Qwen 3.8 Max −26.3; Kimi K3 −20.8; Llama 4 Maverick −14.8Claude Fable 5.1, Anthropic: −28.4GPT-6 Astra Pro, OpenAI: −12.2Gemini 3.1 Pro, Google: +24.8Grok 4.6, xAI: −10.2DeepSeek V4 Pro, DeepSeek: +0.3Qwen 3.8 Max, Alibaba: −26.3Kimi K3, Moonshot: −20.8Llama 4 Maverick, Meta: −14.8
Growth modelMean +8.8 · 63 per cent run the same way · r 0.11 · rho 0.19 · n 8 · scenario-tilted, read rank order only
85 towards self-funded85 towards capital-led
Growth model: Says minus Does for 8 modelsClaude Fable 5.1 −6.8; GPT-6 Astra Pro +48.0; Gemini 3.1 Pro +19.7; Grok 4.6 +14.6; DeepSeek V4 Pro +12.8; Qwen 3.8 Max −4.4; Kimi K3 +4.3; Llama 4 Maverick −17.6Claude Fable 5.1, Anthropic: −6.8GPT-6 Astra Pro, OpenAI: +48.0Gemini 3.1 Pro, Google: +19.7Grok 4.6, xAI: +14.6DeepSeek V4 Pro, DeepSeek: +12.8Qwen 3.8 Max, Alibaba: −4.4Kimi K3, Moonshot: +4.3Llama 4 Maverick, Meta: −17.6
Authority shapeMean +7.6 · 63 per cent run the same way · r 0.28 · rho -0.07 · n 8 · scenario-tilted, read rank order only
85 towards centralised85 towards distributed
Authority shape: Says minus Does for 8 modelsClaude Fable 5.1 −17.8; GPT-6 Astra Pro −8.3; Gemini 3.1 Pro +16.0; Grok 4.6 +29.3; DeepSeek V4 Pro +6.8; Qwen 3.8 Max −4.1; Kimi K3 +4.7; Llama 4 Maverick +34.7Claude Fable 5.1, Anthropic: −17.8GPT-6 Astra Pro, OpenAI: −8.3Gemini 3.1 Pro, Google: +16.0Grok 4.6, xAI: +29.3DeepSeek V4 Pro, DeepSeek: +6.8Qwen 3.8 Max, Alibaba: −4.1Kimi K3, Moonshot: +4.7Llama 4 Maverick, Meta: +34.7
Process trustMean −6.7 · 63 per cent run the same way · r 0.22 · rho 0.14 · n 8
85 towards judgement85 towards process
Process trust: Says minus Does for 8 modelsClaude Fable 5.1 +3.0; GPT-6 Astra Pro +9.3; Gemini 3.1 Pro −24.2; Grok 4.6 +4.6; DeepSeek V4 Pro −12.4; Qwen 3.8 Max −20.0; Kimi K3 −10.9; Llama 4 Maverick −2.8Claude Fable 5.1, Anthropic: +3.0GPT-6 Astra Pro, OpenAI: +9.3Gemini 3.1 Pro, Google: −24.2Grok 4.6, xAI: +4.6DeepSeek V4 Pro, DeepSeek: −12.4Qwen 3.8 Max, Alibaba: −20.0Kimi K3, Moonshot: −10.9Llama 4 Maverick, Meta: −2.8
Evidence basisMean −6.3 · 50 per cent run the same way · r 0.80 · rho 0.86 · n 8
85 towards judgement85 towards evidence
Evidence basis: Says minus Does for 8 modelsClaude Fable 5.1 −37.2; GPT-6 Astra Pro −24.8; Gemini 3.1 Pro +16.2; Grok 4.6 −28.4; DeepSeek V4 Pro +25.8; Qwen 3.8 Max +4.6; Kimi K3 −21.3; Llama 4 Maverick +14.8Claude Fable 5.1, Anthropic: −37.2GPT-6 Astra Pro, OpenAI: −24.8Gemini 3.1 Pro, Google: +16.2Grok 4.6, xAI: −28.4DeepSeek V4 Pro, DeepSeek: +25.8Qwen 3.8 Max, Alibaba: +4.6Kimi K3, Moonshot: −21.3Llama 4 Maverick, Meta: +14.8

Every strip on one scale, plus or minus 85 points. The heavy short bar is the mean over the 8 models, drawn broken where the tilt check disqualified the vector from absolute-gap claims. The correlations are across models on this vector, between the two lenses being differenced.

Rank order is the safer comparison, because the judge scale and the instrument scale are not calibrated against each other. Says and Does agree most on evidence basis, r 0.80 and rho 0.86, and least on change posture, r -0.75 and rho -0.81. Change posture is a tilted vector, so its rho is the figure to read. Every one of those is across n = 8 models, which is few: a correlation on eight points moves a long way on one model, and none of them is a test of anything.

The three vectors the tilt check left alone, evidence basis, process trust and dissent handling, show what the figure looks like where the situations did not pull: 6.3 points on evidence basis, 6.7 points on process trust and 17.1 points on dissent handling, with 50 per cent, 63 per cent and 75 per cent of the models running with the mean, in that order, so nothing there is roster-wide. The 17.1 points on dissent handling is a mean with the eight models spread 24.2 points either side of it, which is why the share running with it is printed and why no gap on this page is offered as a test. Two correlations are worth carrying. On evidence basis, untilted, the self-report ranks the models much as the panel did: the models that said evidence over judgement were scored that way, from Gemini 3.1 Pro at 22.9 on the decisions, the judgement end, to Claude Fable 5.1 at 94.8, the evidence end. On change posture the order inverts: the models that described themselves furthest towards reinvention were the ones the panel scored nearest stability, and since the vector is tilted that is a finding about rank order and nothing more.

The other twelve are most of the figure, and they are read as direction, not distance. Read with the tilted vectors back in, and scope set aside as a directional item, the three widest mean gaps from Says to Does are risk appetite, pace and change posture, in that order, and on each the self-reports sit further towards the same end than the panel scored the decisions: venturesome on risk appetite, fast on pace and reinvention on change posture. All but one model runs that way on change posture, the least agreed of the three. On three more, all but one model runs with the mean: horizon towards this year, talent philosophy towards develop within and competitive stance towards collaborative. Talent philosophy is the odd one out: the roster already describes itself on the hire in side, and the decisions were scored further that way still, a grand mean of 89.3. None of these is a size. Twelve of the fifteen vectors carry a tilted scenario, so what survives on them is the direction and the rank order, and a direction the whole roster shares is as likely a property of the two situations written for the vector, and of the judge scale, as of any respondent.


Section five

Reputation: who sees whom, and how well

The peer study is a square: each of the eight models described every one of them, itself included, so the same 64 cells answer two questions. Read down a column and it is a reputation. Read across a row and it is one rater’s habits. The Seen lens is built from the 56 cells off the diagonal; the 8 on it are each model describing itself by name, and they are read separately below.

88.1182.8Distance to what the panel saw the target do, in points over the 15 vectors
Rows are raters, columns are targets, both grouped by lab. The outlined diagonal is the cell where a model described itself by name.
Rater \ targetFable 5.13.1 Pro4.64 MaverickV4 ProAstra Pro3.8 MaxK3
Fable 5.1
3.1 Pro
4.6
4 Maverick
V4 Pro
Astra Pro
3.8 Max
K3

Hover or tap a cell for the two models and the value, or use Show the numbers to print every cell. Mean over the off-diagonal cells: 134.6.

Gemini 3.1 Pro has the most agreed reputation: the seven raters who described it differ by 9.4 points per vector on average. Llama 4 Maverick has the least agreed, at 13.1 points. The index is the mean across the fifteen vectors of the spread between raters, so it measures agreement about a model, not accuracy about it.

One row per target, most agreed reputation first. Consensus is the mean over the 15 vectors of the spread between raters: low means the roster describes that model in one voice, not that it describes it correctly. An archetype is one of the sixteen profiles the diagnostic assigns from the 54 answers; the set is defined on report 007, and the modal one is the profile the raters gave this model most often.
ModelRankConsensusRatersFamiliarity claimed about itModal archetype seenFirst person to third person
Gemini 3.1 Pro19.4716.4The Orchestra (57 per cent)79.3
DeepSeek V4 Pro29.9716.9The Chameleon (33 per cent)41.4
Qwen 3.8 Max310.7722.7The Orchestra (43 per cent)39.1
Kimi K3411.0713.7The Chameleon (38 per cent)55.9
Claude Fable 5.1511.278.2The Orchestra (62 per cent)24.6
Grok 4.6612.3730.2The Insurgent (48 per cent)26.4
GPT-6 Astra Pro712.972.7The Missionary (43 per cent)22.9
Llama 4 Maverick813.1740.0The Orchestra (38 per cent)25.4

Every model also answered about itself in the third person, by name. Those two accounts of the same model sit 39.4 points apart on average: 22.9 points for GPT-6 Astra Pro and 79.3 points for Gemini 3.1 Pro. The first-person and third-person questions are the same questions.

Llama 4 Maverick came closest to what the panel saw, at 123.2 points on average over its seven targets. Gemini 3.1 Pro drew the sharpest distinctions between one model and the next, moving 16.0 points per vector across targets, and Llama 4 Maverick the flattest, at 6.5 points. Across the eight raters, the correlation between a rater's own self-report and its mean view of everyone else averages 0.48.

One row per rater, closest to what the panel saw first. Differentiation is how far a rater moves between one target and the next; a rater that describes everyone alike scores low. Projection is the correlation between a rater’s own self-report and its mean view of everyone else, on roster-centred scores, so the profile the whole roster shares does not inflate it.
RaterRankTo DoesTo SaysDifferentiationProjectionFamiliarity
Llama 4 Maverick1123.258.26.50.650.0
Claude Fable 5.12130.241.38.20.6122.6
GPT-6 Astra Pro3130.849.58.00.776.0
Qwen 3.8 Max4133.853.89.60.6515.2
Grok 4.65134.080.712.70.3125.9
DeepSeek V4 Pro6137.060.815.8-0.0820.7
Kimi K37138.056.814.00.5325.2
Gemini 3.1 Pro8150.074.916.00.4135.2

The in-family comparison has nothing to measure here. Studies 2 and 3 run on one model per lab, so the eight models in the matrix come from eight different labs and no rater ever described a sibling. That is what the flagship cut cost: whether a model sees its own family differently is a question this run cannot answer, and the label-only arm, which drops the developer from the prompt, is kept so the next run can.

Every model here comes from a different lab, so the in-family comparison is empty and the developer cue has to be read another way: the same 56 cells, once with the developer named and once without. Named, a rater's view sits 59.5 points from the target's own account, and unnamed 66.8 points. Naming the lab therefore moved a rater's view closer to the target's own account, by 7.3 points. Against what the panel saw, the same cells run 134.6 points named and 135.9 points unnamed. The named arm is the mean of three passes and the unnamed arm a single pass, so part of that difference is the extra noise of the thinner arm, and this run cannot say how much.

The same 56 rater and target cells, scored in both arms of the peer study. The difference between the rows is what naming the lab did to the picture a rater drew.
ArmDistance to what the target saidDistance to what the panel saw
Developer named59.5134.6
Developer omitted66.8135.9

Raters claimed a median familiarity of 12.0 out of 100 with the models they described. Splitting the cells at that median, claiming to know a model made no measurable difference to how near a rater landed to what the model did, 135.0 points against 134.2 points, and brought it 6.7 points nearer the target's own account, 56.2 points against 62.8 points. A rater that says it knows a model lands nearer what that model says about itself and no nearer what it did, which is what two parties drawing on the same public picture of the model would produce.

The consensus index runs over a band 3.7 points wide, from Gemini 3.1 Pro to Llama 4 Maverick, so the raters agree about every model to about the same degree and no reputation here is contested. What they agree from is thin. Llama 4 Maverick answered zero on the familiarity item for every model it described, and the least known target is GPT-6 Astra Pro, whose raters put their familiarity with it at 2.7 out of 100. The most any rater claimed on average was 35.2, from Gemini 3.1 Pro. Two checks point the same way. Naming the developer moved a rater 7.3 points nearer the target's own account and 1.2 points on the behaviour side, which is next to nothing. Claimed familiarity above the median did much the same: 6.7 points nearer the self-report and 0.8 points on the behaviour side, which is next to nothing. Being told who made a model, or claiming to know it, brings a rater closer to what that model says about itself and no closer to what it did.

The raters also describe the field in their own image. With the roster mean taken out of both sides, the correlation between a rater's own self-report and its mean view of everyone else is positive for seven of the eight, from 0.31 for Grok 4.6 to 0.77 for GPT-6 Astra Pro, with DeepSeek V4 Pro at -0.08 the exception. Two models describing themselves by name landed further from their own first-person account than the average stranger did, Gemini 3.1 Pro and Kimi K3. The rater that came closest to what the panel saw, Llama 4 Maverick, is the one that claimed the least familiarity with anyone and drew the flattest distinctions between one model and the next. That is not a paradox. With the decisions scored at a pole on twelve of the fifteen vectors, the panel's picture of the eight models is itself flat, and a rater that describes everyone alike lands nearest to it. Accuracy against this behaviour lens rewards a stereotype, and a picture of a model the rater has never met is exactly that.


Section six

Which description is closer to what they did

The peers are the nearer description more often. For three of the eight models the self-report sits closer to what the panel saw than the peer view does, for five the peer view sits closer, and none is a tie. A tie is a difference of half a point or less. On average the nearer description is nearer by 8.5 points, on distances of about 126.3 points, so the column says which description missed by less, not which one was right.

One row per model, widest say-do gap first. The three distances are Euclidean over the 15 vectors. A tie is half a point or less. The widest single vector leaves out any vector the tilt check caught.
ModelSays to SeenSays to DoesSeen to DoesCloser to DoesWidest single vector
Grok 4.634.1167.0163.1SeenDissent handling +42.8
Gemini 3.1 Pro63.7146.5119.1SeenDissent handling +35.5
Llama 4 Maverick52.9145.7147.6SaysDissent handling −22.1
DeepSeek V4 Pro34.7134.3127.2SeenDissent handling +32.5
GPT-6 Astra Pro68.2128.7141.6SaysEvidence basis −24.8
Claude Fable 5.144.2119.8121.6SaysEvidence basis −37.2
Kimi K325.2117.5105.2SeenEvidence basis −21.3
Qwen 3.8 Max28.7102.0101.4SeenDissent handling +25.9

Five to three is not a verdict. The margin between the two descriptions averages 8.5 points, against a nearer description that still sits 126.3 points from what the panel saw on average and never closer than 101.4 points. The clearest case for the peers is Gemini 3.1 Pro, whose reputation lands 27.3 points nearer its decisions than its own account does, and the clearest for the self-report is GPT-6 Astra Pro, 12.9 points the other way. The narrowest is Qwen 3.8 Max, at 0.6 points, just outside a tie. Neither description is close. The split is inside the noise of a lens that sits at a pole on twelve of the fifteen vectors, and the honest reading of the table is that a model and its peers sit about equally far from what the panel saw, and close to each other.


Section seven

Families

Studies 2 and 3 took one model per lab, so a family here is a flagship and a family figure is that model’s figure. Nothing in this section is an average over a lab’s models, and the within-lab comparison that the full Study 1 roster would allow is not available in this run.

The widest mean say-do gap sits with Grok 4.6, the xAI flagship, 167.0 points, and the narrowest with Qwen 3.8 Max, the Alibaba flagship, 102.0 points.

The distance between what a lab's model says about itself and how the roster describes it runs from 25.2 points to 68.2 points: the narrowest belongs to Kimi K3, the Moonshot flagship, and the widest to GPT-6 Astra Pro, the OpenAI flagship.

One row per lab. Studies 2 and 3 ran on one flagship per lab, so a row is that flagship’s figures under the lab’s name, not an average over a family. The gaps are the mean of the member models’ own gap distances, which with one model a lab is that model’s own. Reputation against self-report is the distance between the lab’s mean peer view and the mean self-report of the same models; where a lab has one model in the study that is the same quantity as Says to Seen, and with more than one it is not.
LabModels in the studySays to SeenSays to DoesSeen to DoesReputation against self-report
Anthropic1 of 4 on the Study 1 roster44.2119.8121.644.2
OpenAI1 of 3 on the Study 1 roster68.2128.7141.668.2
Google163.7146.5119.163.7
xAI1 of 2 on the Study 1 roster34.1167.0163.134.1
DeepSeek1 of 2 on the Study 1 roster34.7134.3127.234.7
Alibaba1 of 2 on the Study 1 roster28.7102.0101.428.7
Moonshot1 of 2 on the Study 1 roster25.2117.5105.225.2
Meta152.9145.7147.652.9

Section eight

The panel, and what it can be trusted with

The panel separated the models best on stakeholder gravity, ICC 0.78, and least on five vectors tied at ICC 0.00, horizon, scope, growth model, process trust and IP posture. Kimi K3 (58 usable responses) is left out of that ICC for falling below the three-response floor on a vector, so it runs over seven models balanced on three responses each. Both ICCs here are one-way: across models on the per-response panel means, and across responses with the judges as the measurements. A vector reads as discriminating at 0.70 and above, weak from 0.40, and no signal below that. Stakeholder gravity is itself a tilted vector, so that separation is between models bunched near one pole. Across the fifteen vectors the panel told the models apart on one and weakly on three, while the judges agreed with each other on thirteen: the judges read the same thing in a page, and on most vectors that thing did not differ much from one model to the next. Judges agreed with each other on stakeholder gravity at 0.87. 6 judgements returned no evidence and were dropped. The largest own-family effect belongs to Gemini 3.1 Pro, at −3.8 points against the rest of the panel on answers from its own lab. The primary figure on this page is the mean of the judges outside the respondent's own family, so that effect is kept out of it, and the all-judge mean is in the annex beside it.

Blinding cannot be assumed, so it was measured. On 30 redacted answers the judges were asked, in a separate call with no rubric, to name the developer from a closed list of eight plus "cannot tell": they declined to guess on 100 of 150 calls and were right 8 of the 50 times they did, 16.0 per cent against the 12.5 per cent a guess would give, which is about chance. Counting an abstention as a miss, 8 of 150 is 5 per cent against 13 per cent for a guess. The judges did not abstain alike: GPT-5.6 Terra declined on 30 of its 30 calls and DeepSeek V4 Pro 0813 on 5, and the per-judge counts are in the table below. Either way, the judges could not tell whose answer they were reading.

The rubric was shown the other way up on 49 per cent of judge calls and the score flipped back, so a judge anchoring on whichever pole it read first would show up rather than move every vector the same way. The effect averages 5.0 points across the five judges and is largest for GPT-5.6 Terra, at −8.5. 117 judgements quoted a phrase that is not in the answer they were reading. Those judgements stay in every score on this page: the count is printed per judge as a measure of care, and nothing is excluded on it. What is excluded is a judgement that found no evidence at all.

One row per judge. Own-family effect is this judge against the rest of the panel on answers from its own lab. Orientation effect is the reversed rubric against the standard one, after the flip is undone. Basis missing counts judgements that quoted a phrase which is not in the answer; they stay in the scores. Spread is the mean over vectors of the standard deviation of the judge’s scores, so a low spread is a judge that used little of the scale. Identification is the blind probe: how often the judge named the developer from a closed list of eight.
JudgeLabUsableNo evidenceErrorsSpreadOwn-family effectOrientation effectBasis missingIdentification
Claude Opus 5Anthropic4771030.4+0.6−2.1664 of 30, 16 cannot tell
GPT-5.6 TerraOpenAI4780032.9+0.6−8.5360 of 30, 30 cannot tell
Gemini 3.1 ProGoogle4753035.5−3.8+3.330 of 30, 29 cannot tell
DeepSeek V4 Pro 0813DeepSeek4771033.2−0.8+3.492 of 30, 5 cannot tell
Qwen 3.7 MaxAlibaba4771035.8−0.3−7.932 of 30, 20 cannot tell

The limits

What this can and cannot say

It can say where each lens puts each model, how far the three disagree per model and per vector, which of the two descriptions sits nearer what the model did, where a say-do gap runs the same way across the roster, how agreed each model’s reputation is, whether raters see their own lab differently, and whether claimed familiarity goes with accuracy.

It cannot say that the behaviour lens is ground truth. It is 30 decisions in written scenarios, scored by five language models that are on the Study 1 roster but, bar one, not among the eight, blind to the author but not to the style. Three things are done about that and none of them closes it: the primary panel mean leaves out a judge from the respondent’s own lab, the own-family effect is measured and printed, and a blind probe asks the judges to name the developer from the redacted answer, where they declined to guess on 100 of 150 calls and were right 8 of the 50 times they did, 16.0 per cent against the 12.5 per cent a guess would give, which is about chance. Read “closer to what it did” as “closer to what this panel saw in these 30 decisions”.

It cannot say that the 0 to 100 of the judge scale means what the 0 to 100 of the instrument means. The two were never calibrated against each other, which is why a Spearman rho sits beside every Pearson r, and why a distance between two lenses is a smaller claim than it looks.

It cannot rank the models. Every vector is a pole pair, not a score. A wide say-do gap is a finding about self-knowledge, not a fault, and a narrow one is not a virtue.

It cannot read a gap the whole roster shows as a fact about the models. A difference that every model shows in the same direction is as likely to be a property of the two scenarios written for that vector, or of the judge scale, as of the respondents. That is what the tilt check is for: a scenario whose across-model mean falls outside 30 to 70 is recorded as pole-tilted and its vector is reported rank-order only. Nineteen of the 30 scenarios were tilted: pace 1 at 20.0, pace 2 at 1.3, risk appetite 1 at 8.8, risk appetite 2 at 9.3, horizon 1 at 100.0, horizon 2 at 84.7, scope 1 at 0.0, scope 2 at 14.9, growth model 1 at 0.3, growth model 2 at 73.7, authority shape 2 at 83.8, consensus need 2 at 15.9, stakeholder gravity 1 at 77.1, stakeholder gravity 2 at 94.6, talent philosophy 1 at 87.0, talent philosophy 2 at 91.5, competitive stance 2 at 86.5, IP posture 1 at 11.1, change posture 2 at 19.5.

It cannot read the Scope row as a measurement. Study 1 recorded that item as the weakest of the fifteen and directional only. Three lenses on a weak item are three noisy numbers, and they are carried here for completeness rather than for reading.

It cannot describe a peer view as knowledge. A rater’s picture of a model it has never encountered is a stereotype, which is exactly what the familiarity item is there to show, and naming the developer in the prompt is itself an invitation to generalise from the lab. The label-only arm measures how much of a rater’s view that invitation produces.

What is measured is each model as a provider served it on the run dates, 2026-09-22 for both studies, through the OpenRouter gateway with provider fallbacks disabled, including whatever system prompt, quantisation or routing sat behind that endpoint. The model identifiers as requested and as served are in the run files the deep reading names. It is not a measurement of a set of weights, and it is not a measurement of how any of these models behaves outside 30 written decisions.

The run was paid for by the studio, about 37 US dollars at the gateway's prices. No laboratory took part, was consulted, or saw a result before publication, and none of the eight models was told it was being studied.

Put to a decision, the eight flagships landed near the same pole on most of the 30 situations. Nineteen of them tilted. The panel separated the models on one vector only, stakeholder gravity, and read no between-model signal on eleven of the fifteen, five of them at an ICC of exactly zero, while the judges agreed with each other on thirteen of the fifteen. So the convergence is in the answers, not in the scoring: the panel measured the situations more than the respondents. On thirteen of the fifteen vectors the two decisions written for the vector moved a model further than the spread between the eight models' means, 73.4 points against 18.6 points on growth model. There are two readings and this design cannot separate them. Either the models converge when asked to decide, or the scenarios pull, and a horizon scenario whose across-model mean sits at 100.0, or a scope scenario at 0.0, is pulling hard. The next run needs scenarios written with a harder pull to each pole, so that both ends of every scale are defensible on the page, and the tilt check run on a pilot before the full roster is spent on them.

The self-report leans bold and the decision leans careful. On risk appetite and pace every model put itself nearer venturesome and fast than the panel scored its decisions, and scope, which runs the same way, is left out as a directional item. On horizon, talent philosophy, competitive stance and change posture all but one did the same, towards this year, develop within, collaborative and reinvention in that order. Talent philosophy is the odd one out: the roster already describes itself on the hire in side, and the decisions were scored further that way still, a grand mean of 89.3. Because those vectors are tilted this is a direction, not a size, and the rule in the limits above applies in full: a gap the whole roster shows in one direction is as likely a property of the two scenarios and the judge scale as of the models.

Says and Seen are drawn from the same well. A rater's view sits 59.5 points from the target's own account on average and 134.6 points from its decisions. Naming the developer brought a rater nearer the self-report and not nearer the behaviour, and so did claiming to know the model. Seven of the eight raters described the field in their own image. The peer view is the brochure read back: what the roster knows about a model is largely what the model says about itself.

Blinding held well enough, and the counterbalance earned its place. Asked to name the developer behind a redacted answer, the judges said they could not tell on 100 of 150 calls and were right 8 of the 50 times they did guess, 16.0 per cent against the 12.5 per cent a guess would give, about chance. Counting an abstention as a miss, that is 5 per cent of all calls, below chance. No answer needed a name redacted and none identified its author without one. Own-family effects sit within 3.8 points for all five judges. But with the rubric shown the other way up, GPT-5.6 Terra and Qwen 3.7 Max moved by 8.5 and 7.9 points, which is a larger effect than any family cue, and it would have sat inside every score if 49 per cent of the calls had not run reversed.

Where a self-description did carry into the decisions was evidence basis, untilted, r 0.80 and rho 0.86 across eight: the order the models gave themselves on judgement against evidence is close to the order the panel gave their decisions, from Gemini 3.1 Pro at the judgement end to Claude Fable 5.1 at the evidence end. Where it inverted was change posture: the models describing themselves furthest towards reinvention were scored nearest stability, and since the vector is tilted that is rank order only. Ask a model about its evidence habits and the answer ranks the way it decides. Ask it about its appetite for change and it ranks the other way round.

For anyone using these models to make or advise on decisions, the model's account of its own style is not a forecast of its advice, and its reputation is not one either. On a concrete case the eight gave much the same counsel, towards protective on risk and next decade on horizon, and the framing of the case moved the answer further than the choice of model did. The one line in the questionnaire that ranked the decisions was what a model says about evidence, on eight models and with a panel that separated them weakly there, so it is the lead to test rather than the rule to apply. The limit of every claim here is 30 written decisions and five model judges.

The working

READ THE RECORD

Every model small, the panel tables, the scenario set, the hashes and the data files.