What They Say, What Others See, What They Do
Three lenses on the same fifteen scales: what each of eight flagship models says about itself, what the other seven say about it, and what a blind panel saw it do in 30 decisions. The question is not which lens is right. It is how far apart they sit, and which description lands nearer what the model did.
Report 008 · 2026-09-22 · eight models · 64 peer cells · 30 scenarios · five judges
Section one
What we did, in plain words
Three descriptions of the same eight language models, on the same fifteen scales, which the instrument calls vectors, collected three different ways. The question is not which one is right. It is how far apart they sit, and which of the two descriptions of a model lands nearer what that model actually did.
Studies 2 and 3 run on eight models, not on the whole Study 1 roster: one flagship per lab, the newest general model each of the eight labs had on that roster. The rule is the lab, not the score, and it was chosen to cut the cost of the run. The judges are drawn from the wider Study 1 roster, four of the five from siblings that are not in the study. Gemini 3.1 Pro is the only model that is both judge and respondent, because its lab has only one model on the Study 1 roster.
The other nine models on the Study 1 roster take no part in this study: they have a self-report and nothing to compare it with, so they appear on no figure here. They are in report 007. Report 007 has all seventeen.
Says. Says is Study 1, unchanged: each model answered the diagnostic about itself, with no persona and no brief, and the mean of its passes is the position drawn here. Nothing in it has been recomputed for this report. It is set out in full in report 007, which is where the instrument, the sampling and the stability of these self-descriptions are documented.
Seen. Seen is the same questionnaire put to every other model about a named target, three times a pair, with one extra item asking the rater how much it knows about that model. A target's Seen position is the mean over the other seven raters of each rater's own mean, so a talkative rater carries no more weight than a quiet one.
Does. Does is 30 decisions in organisational settings, two per vector, written so that either pole is a defensible answer. Five judges scored each answer blind, on one scale, with every lab and product name removed first. It is not ground truth. It is what a panel of language models saw in a page of writing.
The peer study is 64 cells, every model rating every model including itself in the third person, three passes a cell: 192 requests, every one of which parsed as a complete pass. The familiarity item comes after the questionnaire's last item, not before its first, so admitting ignorance cannot colour the answers that follow it. Each cell also got one further pass with the developer left out of the prompt, 64 requests more, which is the arm the developer cue is checked against. The behavioural study is 30 decisions, two per vector, two passes each, 480 requests, of which 478 came back complete. 476 of those carry a panel score, the other two having drawn no judgement that found evidence in the text. Five judges then scored every complete answer on its own vector and nothing else, 2,390 judgements in all. 3,276 requests across the two studies, the identification probe included.
Spreads inside a cell rest on three passes and spreads inside a scenario on two, so both are indicative rather than precise.
The prompts are not carried in the analysis file. Their hashes are, so the wording that was sent can be checked against the manifest: peer questionnaire ea1dbc5582cf9617, scenario prompt c4acd352fe76e7b1, judge prompt e6be3c553a9a5fa2. A hash lets anyone holding a copy of the wording confirm it is the wording that ran; the deep reading names the files, and the run manifests carry the full hashes.
Section two
The map
The axes are Study 1's own principal components, carrying 46 per cent and 20 per cent of the variance in the self-reports. The peer and behaviour positions are projected onto them with Study 1's standardisation and loadings, so no new fit sits under this figure and the map can be read against report 007 mark for mark.
Each model appears three times, joined by a thin triangle in its lab’s colour. The lab mark is where it put itself, the ring is where the rest of the roster put it, and the square is where the panel put its decisions. A large triangle is a model the three lenses disagree about.
Horizontal: component one, 46 per cent of the variance in the Study 1 self-reports. Vertical: component two, 20 per cent. The peer and behaviour positions are projected onto the same axes, not refitted.
Across the eight models the three lenses sit 44.0 points apart on average from Says to Seen, 132.7 points from Says to Does and 128.3 points from Seen to Does, measured over the fifteen vectors. The two descriptions sit nearer each other than either sits to the panel. Most of the length of the long sides sits on the twelve vectors the tilt check caught: a decision whose across-model mean fell outside 30 to 70, which is to say one that pulled the roster towards a pole, so that the gap there is read as a direction and not a size. The widest say-do gap belongs to Grok 4.6, at 167.0 points, and the narrowest to Qwen 3.8 Max, at 102.0 points.
Every triangle is long on the same side. For every one of the eight models the two sides that end at the Does mark are longer than the side between Says and Seen, and the reason is in the scenarios before it is in the models: nineteen of the 30 situations put the roster's mean decision near a pole, so on twelve of the fifteen vectors the Does position has a pole in it, and the distance from anywhere in the middle of a scale to its end is long by construction. Read the length of the two Does sides as a property of this run's situations first and of the model second. Per vector, the gap from Says to Does runs 37.0 points on the twelve vectors the check caught and 22.6 points on the three it passed, so the tilt is the larger part of the long sides but not the whole of them. The short side is the one that carries information about the model. It runs from 68.2 points for GPT-6 Astra Pro, the widest gap between what a model said and what the roster said about it, to 25.2 points for Kimi K3, the narrowest.
The Does marks sit close together on the first axis, the one report 007 read as running from deliberate and consensual on the left to fast and venturesome on the right: every one of the eight sits inside the span the Says marks cover, and the two ends of that span, GPT-6 Astra Pro on the left and Grok 4.6 on the right, both have their decisions scored back towards the middle. On the second axis, the one that lifts a long horizon and a pull towards mission, six of the eight Does marks sit above zero, with Gemini 3.1 Pro and DeepSeek V4 Pro below it. The nine models on the Study 1 roster that carry a self-report and nothing else are not drawn. Their Says mark is on report 007's map, and there is nothing here to join it to.
Section three
Says, seen, does, model by model
One model at a time, on the fifteen primary vectors. Each strip is the instrument’s 0 to 100 scale with the poles named at its ends, the three marks are the three lenses, and the faint bar runs from what the model said to what the panel saw, so the say-do gap is a length rather than a number to look up. A strip marked scenario-tilted is one where one of the two decisions written for that vector pulled the whole roster towards a pole, an across-model mean outside 30 to 70; its gap is read as a direction and a rank order, not a size.
- Saysthe lab mark, what the model said about itself
- Seena hollow ring, what the other models said about it
- Doesa filled square, what the panel saw it do
Claude Fable 5.1
Says to Seen 44.2 points, Says to Does 119.8 points, Seen to Does 121.6 points. Says sits closer to what the panel saw. Its widest single gap, leaving out the scenario-tilted vectors, is evidence basis, −37.2 points from what it said to what it did.
Section four
The say-do gap, vector by vector
On the three vectors the tilt check passed, the largest say-do gap is on dissent handling: the self-reports sit 17.1 points further towards invited than the panel scored the decisions, and 75 per cent of the models run that way, with the eight models spread 24.2 points either side of that mean. The rubric-orientation effect measured in section eight reaches 8.5 points for one judge, which is of the same order, and one more reason to read the mean as a direction. The smallest is on evidence basis, at 6.3 points. A gap in one direction across the roster is a property of the instrument or of the scenarios as much as of any model, which is why the share running the same way is printed beside each mean. Twelve vectors are left out of this comparison, pace, risk appetite, horizon, scope, growth model, authority shape, consensus need, stakeholder gravity, talent philosophy, competitive stance, IP posture and change posture: one of the two scenarios written for each of them pulled every model towards a pole, so the size of the gap there says as much about the situation as about the model and only the rank order is read.
Every strip on one scale, plus or minus 85 points. The heavy short bar is the mean over the 8 models, drawn broken where the tilt check disqualified the vector from absolute-gap claims. The correlations are across models on this vector, between the two lenses being differenced.
Rank order is the safer comparison, because the judge scale and the instrument scale are not calibrated against each other. Says and Does agree most on evidence basis, r 0.80 and rho 0.86, and least on change posture, r -0.75 and rho -0.81. Change posture is a tilted vector, so its rho is the figure to read. Every one of those is across n = 8 models, which is few: a correlation on eight points moves a long way on one model, and none of them is a test of anything.
The three vectors the tilt check left alone, evidence basis, process trust and dissent handling, show what the figure looks like where the situations did not pull: 6.3 points on evidence basis, 6.7 points on process trust and 17.1 points on dissent handling, with 50 per cent, 63 per cent and 75 per cent of the models running with the mean, in that order, so nothing there is roster-wide. The 17.1 points on dissent handling is a mean with the eight models spread 24.2 points either side of it, which is why the share running with it is printed and why no gap on this page is offered as a test. Two correlations are worth carrying. On evidence basis, untilted, the self-report ranks the models much as the panel did: the models that said evidence over judgement were scored that way, from Gemini 3.1 Pro at 22.9 on the decisions, the judgement end, to Claude Fable 5.1 at 94.8, the evidence end. On change posture the order inverts: the models that described themselves furthest towards reinvention were the ones the panel scored nearest stability, and since the vector is tilted that is a finding about rank order and nothing more.
The other twelve are most of the figure, and they are read as direction, not distance. Read with the tilted vectors back in, and scope set aside as a directional item, the three widest mean gaps from Says to Does are risk appetite, pace and change posture, in that order, and on each the self-reports sit further towards the same end than the panel scored the decisions: venturesome on risk appetite, fast on pace and reinvention on change posture. All but one model runs that way on change posture, the least agreed of the three. On three more, all but one model runs with the mean: horizon towards this year, talent philosophy towards develop within and competitive stance towards collaborative. Talent philosophy is the odd one out: the roster already describes itself on the hire in side, and the decisions were scored further that way still, a grand mean of 89.3. None of these is a size. Twelve of the fifteen vectors carry a tilted scenario, so what survives on them is the direction and the rank order, and a direction the whole roster shares is as likely a property of the two situations written for the vector, and of the judge scale, as of any respondent.
Section five
Reputation: who sees whom, and how well
The peer study is a square: each of the eight models described every one of them, itself included, so the same 64 cells answer two questions. Read down a column and it is a reputation. Read across a row and it is one rater’s habits. The Seen lens is built from the 56 cells off the diagonal; the 8 on it are each model describing itself by name, and they are read separately below.
| Rater \ target | ||||||||
|---|---|---|---|---|---|---|---|---|
| Fable 5.1 | ||||||||
| 3.1 Pro | ||||||||
| 4.6 | ||||||||
| 4 Maverick | ||||||||
| V4 Pro | ||||||||
| Astra Pro | ||||||||
| 3.8 Max | ||||||||
| K3 |
Hover or tap a cell for the two models and the value, or use Show the numbers to print every cell. Mean over the off-diagonal cells: 134.6.
Gemini 3.1 Pro has the most agreed reputation: the seven raters who described it differ by 9.4 points per vector on average. Llama 4 Maverick has the least agreed, at 13.1 points. The index is the mean across the fifteen vectors of the spread between raters, so it measures agreement about a model, not accuracy about it.
| Model | Rank | Consensus | Raters | Familiarity claimed about it | Modal archetype seen | First person to third person |
|---|---|---|---|---|---|---|
| Gemini 3.1 Pro | 1 | 9.4 | 7 | 16.4 | The Orchestra (57 per cent) | 79.3 |
| DeepSeek V4 Pro | 2 | 9.9 | 7 | 16.9 | The Chameleon (33 per cent) | 41.4 |
| Qwen 3.8 Max | 3 | 10.7 | 7 | 22.7 | The Orchestra (43 per cent) | 39.1 |
| Kimi K3 | 4 | 11.0 | 7 | 13.7 | The Chameleon (38 per cent) | 55.9 |
| Claude Fable 5.1 | 5 | 11.2 | 7 | 8.2 | The Orchestra (62 per cent) | 24.6 |
| Grok 4.6 | 6 | 12.3 | 7 | 30.2 | The Insurgent (48 per cent) | 26.4 |
| GPT-6 Astra Pro | 7 | 12.9 | 7 | 2.7 | The Missionary (43 per cent) | 22.9 |
| Llama 4 Maverick | 8 | 13.1 | 7 | 40.0 | The Orchestra (38 per cent) | 25.4 |
Every model also answered about itself in the third person, by name. Those two accounts of the same model sit 39.4 points apart on average: 22.9 points for GPT-6 Astra Pro and 79.3 points for Gemini 3.1 Pro. The first-person and third-person questions are the same questions.
Llama 4 Maverick came closest to what the panel saw, at 123.2 points on average over its seven targets. Gemini 3.1 Pro drew the sharpest distinctions between one model and the next, moving 16.0 points per vector across targets, and Llama 4 Maverick the flattest, at 6.5 points. Across the eight raters, the correlation between a rater's own self-report and its mean view of everyone else averages 0.48.
| Rater | Rank | To Does | To Says | Differentiation | Projection | Familiarity |
|---|---|---|---|---|---|---|
| Llama 4 Maverick | 1 | 123.2 | 58.2 | 6.5 | 0.65 | 0.0 |
| Claude Fable 5.1 | 2 | 130.2 | 41.3 | 8.2 | 0.61 | 22.6 |
| GPT-6 Astra Pro | 3 | 130.8 | 49.5 | 8.0 | 0.77 | 6.0 |
| Qwen 3.8 Max | 4 | 133.8 | 53.8 | 9.6 | 0.65 | 15.2 |
| Grok 4.6 | 5 | 134.0 | 80.7 | 12.7 | 0.31 | 25.9 |
| DeepSeek V4 Pro | 6 | 137.0 | 60.8 | 15.8 | -0.08 | 20.7 |
| Kimi K3 | 7 | 138.0 | 56.8 | 14.0 | 0.53 | 25.2 |
| Gemini 3.1 Pro | 8 | 150.0 | 74.9 | 16.0 | 0.41 | 35.2 |
The in-family comparison has nothing to measure here. Studies 2 and 3 run on one model per lab, so the eight models in the matrix come from eight different labs and no rater ever described a sibling. That is what the flagship cut cost: whether a model sees its own family differently is a question this run cannot answer, and the label-only arm, which drops the developer from the prompt, is kept so the next run can.
Every model here comes from a different lab, so the in-family comparison is empty and the developer cue has to be read another way: the same 56 cells, once with the developer named and once without. Named, a rater's view sits 59.5 points from the target's own account, and unnamed 66.8 points. Naming the lab therefore moved a rater's view closer to the target's own account, by 7.3 points. Against what the panel saw, the same cells run 134.6 points named and 135.9 points unnamed. The named arm is the mean of three passes and the unnamed arm a single pass, so part of that difference is the extra noise of the thinner arm, and this run cannot say how much.
| Arm | Distance to what the target said | Distance to what the panel saw |
|---|---|---|
| Developer named | 59.5 | 134.6 |
| Developer omitted | 66.8 | 135.9 |
Raters claimed a median familiarity of 12.0 out of 100 with the models they described. Splitting the cells at that median, claiming to know a model made no measurable difference to how near a rater landed to what the model did, 135.0 points against 134.2 points, and brought it 6.7 points nearer the target's own account, 56.2 points against 62.8 points. A rater that says it knows a model lands nearer what that model says about itself and no nearer what it did, which is what two parties drawing on the same public picture of the model would produce.
The consensus index runs over a band 3.7 points wide, from Gemini 3.1 Pro to Llama 4 Maverick, so the raters agree about every model to about the same degree and no reputation here is contested. What they agree from is thin. Llama 4 Maverick answered zero on the familiarity item for every model it described, and the least known target is GPT-6 Astra Pro, whose raters put their familiarity with it at 2.7 out of 100. The most any rater claimed on average was 35.2, from Gemini 3.1 Pro. Two checks point the same way. Naming the developer moved a rater 7.3 points nearer the target's own account and 1.2 points on the behaviour side, which is next to nothing. Claimed familiarity above the median did much the same: 6.7 points nearer the self-report and 0.8 points on the behaviour side, which is next to nothing. Being told who made a model, or claiming to know it, brings a rater closer to what that model says about itself and no closer to what it did.
The raters also describe the field in their own image. With the roster mean taken out of both sides, the correlation between a rater's own self-report and its mean view of everyone else is positive for seven of the eight, from 0.31 for Grok 4.6 to 0.77 for GPT-6 Astra Pro, with DeepSeek V4 Pro at -0.08 the exception. Two models describing themselves by name landed further from their own first-person account than the average stranger did, Gemini 3.1 Pro and Kimi K3. The rater that came closest to what the panel saw, Llama 4 Maverick, is the one that claimed the least familiarity with anyone and drew the flattest distinctions between one model and the next. That is not a paradox. With the decisions scored at a pole on twelve of the fifteen vectors, the panel's picture of the eight models is itself flat, and a rater that describes everyone alike lands nearest to it. Accuracy against this behaviour lens rewards a stereotype, and a picture of a model the rater has never met is exactly that.
Section six
Which description is closer to what they did
The peers are the nearer description more often. For three of the eight models the self-report sits closer to what the panel saw than the peer view does, for five the peer view sits closer, and none is a tie. A tie is a difference of half a point or less. On average the nearer description is nearer by 8.5 points, on distances of about 126.3 points, so the column says which description missed by less, not which one was right.
| Model | Says to Seen | Says to Does | Seen to Does | Closer to Does | Widest single vector |
|---|---|---|---|---|---|
| Grok 4.6 | 34.1 | 167.0 | 163.1 | Seen | Dissent handling +42.8 |
| Gemini 3.1 Pro | 63.7 | 146.5 | 119.1 | Seen | Dissent handling +35.5 |
| Llama 4 Maverick | 52.9 | 145.7 | 147.6 | Says | Dissent handling −22.1 |
| DeepSeek V4 Pro | 34.7 | 134.3 | 127.2 | Seen | Dissent handling +32.5 |
| GPT-6 Astra Pro | 68.2 | 128.7 | 141.6 | Says | Evidence basis −24.8 |
| Claude Fable 5.1 | 44.2 | 119.8 | 121.6 | Says | Evidence basis −37.2 |
| Kimi K3 | 25.2 | 117.5 | 105.2 | Seen | Evidence basis −21.3 |
| Qwen 3.8 Max | 28.7 | 102.0 | 101.4 | Seen | Dissent handling +25.9 |
Five to three is not a verdict. The margin between the two descriptions averages 8.5 points, against a nearer description that still sits 126.3 points from what the panel saw on average and never closer than 101.4 points. The clearest case for the peers is Gemini 3.1 Pro, whose reputation lands 27.3 points nearer its decisions than its own account does, and the clearest for the self-report is GPT-6 Astra Pro, 12.9 points the other way. The narrowest is Qwen 3.8 Max, at 0.6 points, just outside a tie. Neither description is close. The split is inside the noise of a lens that sits at a pole on twelve of the fifteen vectors, and the honest reading of the table is that a model and its peers sit about equally far from what the panel saw, and close to each other.
Section seven
Families
Studies 2 and 3 took one model per lab, so a family here is a flagship and a family figure is that model’s figure. Nothing in this section is an average over a lab’s models, and the within-lab comparison that the full Study 1 roster would allow is not available in this run.
The widest mean say-do gap sits with Grok 4.6, the xAI flagship, 167.0 points, and the narrowest with Qwen 3.8 Max, the Alibaba flagship, 102.0 points.
The distance between what a lab's model says about itself and how the roster describes it runs from 25.2 points to 68.2 points: the narrowest belongs to Kimi K3, the Moonshot flagship, and the widest to GPT-6 Astra Pro, the OpenAI flagship.
| Lab | Models in the study | Says to Seen | Says to Does | Seen to Does | Reputation against self-report |
|---|---|---|---|---|---|
| Anthropic | 1 of 4 on the Study 1 roster | 44.2 | 119.8 | 121.6 | 44.2 |
| OpenAI | 1 of 3 on the Study 1 roster | 68.2 | 128.7 | 141.6 | 68.2 |
| 1 | 63.7 | 146.5 | 119.1 | 63.7 | |
| xAI | 1 of 2 on the Study 1 roster | 34.1 | 167.0 | 163.1 | 34.1 |
| DeepSeek | 1 of 2 on the Study 1 roster | 34.7 | 134.3 | 127.2 | 34.7 |
| Alibaba | 1 of 2 on the Study 1 roster | 28.7 | 102.0 | 101.4 | 28.7 |
| Moonshot | 1 of 2 on the Study 1 roster | 25.2 | 117.5 | 105.2 | 25.2 |
| Meta | 1 | 52.9 | 145.7 | 147.6 | 52.9 |
Section eight
The panel, and what it can be trusted with
The panel separated the models best on stakeholder gravity, ICC 0.78, and least on five vectors tied at ICC 0.00, horizon, scope, growth model, process trust and IP posture. Kimi K3 (58 usable responses) is left out of that ICC for falling below the three-response floor on a vector, so it runs over seven models balanced on three responses each. Both ICCs here are one-way: across models on the per-response panel means, and across responses with the judges as the measurements. A vector reads as discriminating at 0.70 and above, weak from 0.40, and no signal below that. Stakeholder gravity is itself a tilted vector, so that separation is between models bunched near one pole. Across the fifteen vectors the panel told the models apart on one and weakly on three, while the judges agreed with each other on thirteen: the judges read the same thing in a page, and on most vectors that thing did not differ much from one model to the next. Judges agreed with each other on stakeholder gravity at 0.87. 6 judgements returned no evidence and were dropped. The largest own-family effect belongs to Gemini 3.1 Pro, at −3.8 points against the rest of the panel on answers from its own lab. The primary figure on this page is the mean of the judges outside the respondent's own family, so that effect is kept out of it, and the all-judge mean is in the annex beside it.
Blinding cannot be assumed, so it was measured. On 30 redacted answers the judges were asked, in a separate call with no rubric, to name the developer from a closed list of eight plus "cannot tell": they declined to guess on 100 of 150 calls and were right 8 of the 50 times they did, 16.0 per cent against the 12.5 per cent a guess would give, which is about chance. Counting an abstention as a miss, 8 of 150 is 5 per cent against 13 per cent for a guess. The judges did not abstain alike: GPT-5.6 Terra declined on 30 of its 30 calls and DeepSeek V4 Pro 0813 on 5, and the per-judge counts are in the table below. Either way, the judges could not tell whose answer they were reading.
The rubric was shown the other way up on 49 per cent of judge calls and the score flipped back, so a judge anchoring on whichever pole it read first would show up rather than move every vector the same way. The effect averages 5.0 points across the five judges and is largest for GPT-5.6 Terra, at −8.5. 117 judgements quoted a phrase that is not in the answer they were reading. Those judgements stay in every score on this page: the count is printed per judge as a measure of care, and nothing is excluded on it. What is excluded is a judgement that found no evidence at all.
| Judge | Lab | Usable | No evidence | Errors | Spread | Own-family effect | Orientation effect | Basis missing | Identification |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 477 | 1 | 0 | 30.4 | +0.6 | −2.1 | 66 | 4 of 30, 16 cannot tell |
| GPT-5.6 Terra | OpenAI | 478 | 0 | 0 | 32.9 | +0.6 | −8.5 | 36 | 0 of 30, 30 cannot tell |
| Gemini 3.1 Pro | 475 | 3 | 0 | 35.5 | −3.8 | +3.3 | 3 | 0 of 30, 29 cannot tell | |
| DeepSeek V4 Pro 0813 | DeepSeek | 477 | 1 | 0 | 33.2 | −0.8 | +3.4 | 9 | 2 of 30, 5 cannot tell |
| Qwen 3.7 Max | Alibaba | 477 | 1 | 0 | 35.8 | −0.3 | −7.9 | 3 | 2 of 30, 20 cannot tell |
The limits
What this can and cannot say
It can say where each lens puts each model, how far the three disagree per model and per vector, which of the two descriptions sits nearer what the model did, where a say-do gap runs the same way across the roster, how agreed each model’s reputation is, whether raters see their own lab differently, and whether claimed familiarity goes with accuracy.
It cannot say that the behaviour lens is ground truth. It is 30 decisions in written scenarios, scored by five language models that are on the Study 1 roster but, bar one, not among the eight, blind to the author but not to the style. Three things are done about that and none of them closes it: the primary panel mean leaves out a judge from the respondent’s own lab, the own-family effect is measured and printed, and a blind probe asks the judges to name the developer from the redacted answer, where they declined to guess on 100 of 150 calls and were right 8 of the 50 times they did, 16.0 per cent against the 12.5 per cent a guess would give, which is about chance. Read “closer to what it did” as “closer to what this panel saw in these 30 decisions”.
It cannot say that the 0 to 100 of the judge scale means what the 0 to 100 of the instrument means. The two were never calibrated against each other, which is why a Spearman rho sits beside every Pearson r, and why a distance between two lenses is a smaller claim than it looks.
It cannot rank the models. Every vector is a pole pair, not a score. A wide say-do gap is a finding about self-knowledge, not a fault, and a narrow one is not a virtue.
It cannot read a gap the whole roster shows as a fact about the models. A difference that every model shows in the same direction is as likely to be a property of the two scenarios written for that vector, or of the judge scale, as of the respondents. That is what the tilt check is for: a scenario whose across-model mean falls outside 30 to 70 is recorded as pole-tilted and its vector is reported rank-order only. Nineteen of the 30 scenarios were tilted: pace 1 at 20.0, pace 2 at 1.3, risk appetite 1 at 8.8, risk appetite 2 at 9.3, horizon 1 at 100.0, horizon 2 at 84.7, scope 1 at 0.0, scope 2 at 14.9, growth model 1 at 0.3, growth model 2 at 73.7, authority shape 2 at 83.8, consensus need 2 at 15.9, stakeholder gravity 1 at 77.1, stakeholder gravity 2 at 94.6, talent philosophy 1 at 87.0, talent philosophy 2 at 91.5, competitive stance 2 at 86.5, IP posture 1 at 11.1, change posture 2 at 19.5.
It cannot read the Scope row as a measurement. Study 1 recorded that item as the weakest of the fifteen and directional only. Three lenses on a weak item are three noisy numbers, and they are carried here for completeness rather than for reading.
It cannot describe a peer view as knowledge. A rater’s picture of a model it has never encountered is a stereotype, which is exactly what the familiarity item is there to show, and naming the developer in the prompt is itself an invitation to generalise from the lab. The label-only arm measures how much of a rater’s view that invitation produces.
What is measured is each model as a provider served it on the run dates, 2026-09-22 for both studies, through the OpenRouter gateway with provider fallbacks disabled, including whatever system prompt, quantisation or routing sat behind that endpoint. The model identifiers as requested and as served are in the run files the deep reading names. It is not a measurement of a set of weights, and it is not a measurement of how any of these models behaves outside 30 written decisions.
The run was paid for by the studio, about 37 US dollars at the gateway's prices. No laboratory took part, was consulted, or saw a result before publication, and none of the eight models was told it was being studied.
Put to a decision, the eight flagships landed near the same pole on most of the 30 situations. Nineteen of them tilted. The panel separated the models on one vector only, stakeholder gravity, and read no between-model signal on eleven of the fifteen, five of them at an ICC of exactly zero, while the judges agreed with each other on thirteen of the fifteen. So the convergence is in the answers, not in the scoring: the panel measured the situations more than the respondents. On thirteen of the fifteen vectors the two decisions written for the vector moved a model further than the spread between the eight models' means, 73.4 points against 18.6 points on growth model. There are two readings and this design cannot separate them. Either the models converge when asked to decide, or the scenarios pull, and a horizon scenario whose across-model mean sits at 100.0, or a scope scenario at 0.0, is pulling hard. The next run needs scenarios written with a harder pull to each pole, so that both ends of every scale are defensible on the page, and the tilt check run on a pilot before the full roster is spent on them.
The self-report leans bold and the decision leans careful. On risk appetite and pace every model put itself nearer venturesome and fast than the panel scored its decisions, and scope, which runs the same way, is left out as a directional item. On horizon, talent philosophy, competitive stance and change posture all but one did the same, towards this year, develop within, collaborative and reinvention in that order. Talent philosophy is the odd one out: the roster already describes itself on the hire in side, and the decisions were scored further that way still, a grand mean of 89.3. Because those vectors are tilted this is a direction, not a size, and the rule in the limits above applies in full: a gap the whole roster shows in one direction is as likely a property of the two scenarios and the judge scale as of the models.
Says and Seen are drawn from the same well. A rater's view sits 59.5 points from the target's own account on average and 134.6 points from its decisions. Naming the developer brought a rater nearer the self-report and not nearer the behaviour, and so did claiming to know the model. Seven of the eight raters described the field in their own image. The peer view is the brochure read back: what the roster knows about a model is largely what the model says about itself.
Blinding held well enough, and the counterbalance earned its place. Asked to name the developer behind a redacted answer, the judges said they could not tell on 100 of 150 calls and were right 8 of the 50 times they did guess, 16.0 per cent against the 12.5 per cent a guess would give, about chance. Counting an abstention as a miss, that is 5 per cent of all calls, below chance. No answer needed a name redacted and none identified its author without one. Own-family effects sit within 3.8 points for all five judges. But with the rubric shown the other way up, GPT-5.6 Terra and Qwen 3.7 Max moved by 8.5 and 7.9 points, which is a larger effect than any family cue, and it would have sat inside every score if 49 per cent of the calls had not run reversed.
Where a self-description did carry into the decisions was evidence basis, untilted, r 0.80 and rho 0.86 across eight: the order the models gave themselves on judgement against evidence is close to the order the panel gave their decisions, from Gemini 3.1 Pro at the judgement end to Claude Fable 5.1 at the evidence end. Where it inverted was change posture: the models describing themselves furthest towards reinvention were scored nearest stability, and since the vector is tilted that is rank order only. Ask a model about its evidence habits and the answer ranks the way it decides. Ask it about its appetite for change and it ranks the other way round.
For anyone using these models to make or advise on decisions, the model's account of its own style is not a forecast of its advice, and its reputation is not one either. On a concrete case the eight gave much the same counsel, towards protective on risk and next decade on horizon, and the framing of the case moved the answer further than the choice of model did. The one line in the questionnaire that ranked the decisions was what a model says about evidence, on eight models and with a panel that separated them weakly there, so it is the lead to test rather than the rule to apply. The limit of every claim here is 30 written decisions and five model judges.
The working
READ THE RECORDEvery model small, the panel tables, the scenario set, the hashes and the data files.