Skip to content

We asked eight language models what they are like, asked the others, then watched what they did.

Three descriptions of the same model, one flagship from each lab, on the same fifteen scales. This is where they disagree.

Report 008 · 2026-09-22 · eight models · 64 peer cells · 30 scenarios · five judges

TLDR

1We measured eight language models three ways, one from each lab: what each says about itself, what the other seven say about it, and what a blind panel saw it do in 30 decisions.

2On twelve of the fifteen scales one of the two decisions pulled the whole roster towards a pole, so a say-do gap there is read as a direction and a rank order, not a size. The three that stayed clear are evidence basis, process trust and dissent handling.

3On the three scales the tilt check passed, the biggest shared gap is on dissent handling: the models describe themselves 17.1 points further towards invited than the panel scored their decisions.

4For three of the eight models the model's own account lands nearer its decisions than its peers' account does, and for five it is the other way around. The margins average 8.5 points over the eight, against distances of about 126.3 points.

5The other seven were mostly describing a model they said they barely knew: the median claimed familiarity was 12.0 out of 100.

6The panel is not ground truth. It is five language models reading 476 short written answers, blind to the author.

7If you have to choose a model for a decision, the finding is to write the decision out and try it: the wording moved these models further than the choice between them did.

How it was made

The peer study is 64 cells, every model rating every model including itself in the third person, three passes a cell: 192 requests, every one of which parsed as a complete pass. The familiarity item comes after the questionnaire's last item, not before its first, so admitting ignorance cannot colour the answers that follow it. Each cell also got one further pass with the developer left out of the prompt, 64 requests more, which is the arm the developer cue is checked against. The behavioural study is 30 decisions, two per vector, two passes each, 480 requests, of which 478 came back complete. 476 of those carry a panel score, the other two having drawn no judgement that found evidence in the text. Five judges then scored every complete answer on its own vector and nothing else, 2,390 judgements in all. 3,276 requests across the two studies, the identification probe included.

The questionnaire is the one this site gives an organisation. The decisions are ordinary organisational problems where either answer is defensible, and the judges never saw who wrote the answer.

The Says lens is the earlier study of the same models, report 007, unchanged. Nothing here was rescored to make the three agree.

The map

Each model is on here three times: the lab mark is what it says about itself, the ring is what the other models say about it, and the square is what the panel saw it do. The thin triangle joins them, so a big triangle is a model the three lenses disagree about.

8 models under three lenses, on Study 1’s two principal componentsScatter plot. Each model appears three times: as its lab’s logo where it placed itself, as a hollow ring where the other models placed it, and as a filled square where the panel placed its decisions. A thin triangle in the family colour joins the three.

Horizontal: component one, 46 per cent of the variance in the Study 1 self-reports. Vertical: component two, 20 per cent. The peer and behaviour positions are projected onto the same axes, not refitted.

Three things to notice

The decision moved them more than who they are did

On thirteen of the fifteen scales the two decisions written for the scale moved a model further than the eight models differ from each other. On growth model the two settings sit 73.4 points apart per model, against 18.6 points between models.

Where saying and doing part company

On dissent handling the self-reports sit 17.1 points from the panel's reading, and 75 per cent of the models lean the same way. On evidence basis the gap is 6.3 points.

Reputation without acquaintance

Raters claimed a median familiarity of 12.0 out of 100 with the models they described, and taking the developer's name out of the prompt moved their views 7.3 points further from the target's own account. Part of a reputation here is the lab on the label.

What it cannot say

It cannot say that the panel is right. The panel is five language models reading 476 short written answers, blind to who wrote them but not to the style.

It does not rank the models. Every scale has two ends, and a wide gap between what a model says and what it did is a fact about self-description, not a fault.

A gap the whole roster shows is as likely to be a property of the scenarios as of the models, which is why twelve of the fifteen scales here are reported by rank order only.

The line to repeat

Eight models that describe themselves differently were scored near one pole on twelve of fifteen scales, and the report cannot say whether that is the models or the decisions. On the three where the situations did not pull, dissent handling is where what they say and what they did part company most.

Got it?

READ NORMAL MODE