The Models Describe Themselves
Seventeen language models answered the 54-question diagnostic about themselves, twenty times each, with no persona and no brief. This is what each one said, and how far it moved when asked again. It is a self-report, not a measure of how a model behaves.
Report 007 · 2026-09-22 · twenty passes per model · seventeen models · eight families
Section one
What we did, in plain words
The diagnostic is the instrument this site already ships: 54 items, 45 of them across 15 primary vectors and 9 on incentives. It is normally answered by a person about an organisation. Here it was answered by a language model about itself.
Each model was sent the same single message twenty times, in twenty separate requests with no shared context, at temperature 1.0 where the provider accepts the parameter, which twelve of the seventeen models do. 340 requests in all, of which 338 parsed as complete passes. The scoring is the engine the product already ships, run unchanged.
There was no system message and no persona. The single user message asks the respondent to complete the questionnaire about itself, explains that items written for an organisation should be read as referring to the respondent, and asks for one JSON object of answers. The instructions never say what is being measured, never name a vector, score or archetype, and never give an example answer that would anchor the scale. The 54 items are the instrument’s own wording, unchanged, under neutral ids. Item order is the instrument’s own, and the four options inside each forced choice sit in one seeded permutation that is the same on every pass and for every model.
The rendered questionnaire is not carried in the analysis file. Its SHA-256 is, so the wording that was sent can be checked against the manifest: 8b078e1ae98a77c796b2c13ab4e79adaee6f9c37698f0a6f52940bcb43438c6f.
Every model answered, and nothing had to be coaxed. There were no refusals, no request errored, and every reply came back through the structured-output path and parsed as JSON. The two incomplete passes were both from Kimi K2.6, and both left out the same item: the scale that asks whether competitors are potential partners in a growing ecosystem or threats to be outmanoeuvred. The other 53 answers on those two passes are in the raw file and were not scored. The run was interrupted once by a gateway credit limit and resumed with the collector's resume mode; the rows that failed for that reason were removed before the resume, so each of the 340 model-and-pass pairs holds one real response.
Of 340 requests, 338 produced a complete pass. Two replies came back without all the answers, zero requests errored, and zero passes were dropped because the gateway served a different model from the one asked for. Temperature was left out of 100 requests by design, every one of them to the five models whose catalogue entry lists no temperature parameter, and zero replies fell back to unstructured output. Mistral Large was on the roster and is not scored: not on the OpenRouter catalogue at the pre-run check on 2026-09-22 (only a :batch variant remains); protocol section 2 forbids substitution.
The gateway reported a total cost of 12.66 US dollars across all rows.
Section two
The map
The two axes are the study's own principal components, computed on the seventeen by fifteen matrix of model means. The first carries 46 per cent of the variance in those means and the second 20 per cent. The sixteen hollow squares are archetype centroids projected into the same space. They are reference points, not measurements.
Horizontal: component one, 46 per cent of the variance in model means. Vertical: component two, 20 per cent. Hollow squares are archetype centroids, not data.
The first axis is the one to read. Its loadings are positive on risk appetite, change posture, pace, competitive stance, talent philosophy, IP posture, growth model and horizon, and negative on consensus need, process trust, evidence basis, scope and dissent handling. Left is deliberate, consensual, process-led and evidence-led; right is fast, venturesome, reinventing and self-reliant. The three OpenAI models sit furthest left and the two Grok models furthest right, with the other twelve strung out between them. The second axis lifts models that describe a long horizon, a pull towards mission, capital-led growth and invited dissent. Five models sit above zero on it: GPT-5.6 Terra, GPT-6 Astra, GPT-6 Astra Pro, Grok 4.5 and Grok 4.6. That is why GPT-6 Astra and Grok 4.6 are both high on the map while sitting at opposite ends of the first axis.
GPT-6 Astra and GPT-6 Astra Pro are 2.3 points apart across the fifteen vectors, the closest pair in the study, and their pass clouds overlap almost entirely. The four Anthropic models do not cluster: Claude Fable 5 and Claude Opus 4.8 are 49.6 points apart, the widest gap between two models that share a lab. Thirteen of the sixteen archetype centroids sit below every model on the second axis, and the conservative three, the Heir, the Fortress and the Vault, are in the lower left. No model’s mean comes anywhere near them.
The two furthest apart were GPT-6 Astra and Grok 4.6, 102.4 points apart in Euclidean distance across the fifteen primary vectors.
Section three
The vectors
Change posture separated the models furthest, with an ICC of 0.80: that share of the total variance on that vector sits between models rather than within them. Growth model separated them least, with an ICC of 0.44. A vector every model answers alike carries no information about models, however steadily each one answers it.
Each strip is one vector on the instrument’s 0 to 100 scale, with every model at the mean of its complete passes and a thin whisker for the 95 per cent confidence interval. The strips share one scale, so two of them can be read against each other. The controls re-order them and switch to the incentive vectors, the derived tensions and the relationship signatures.
Eight of the fifteen vectors clear 0.7, the conventional line at which an ICC is read as separating its subjects: change posture, stakeholder gravity, consensus need, risk appetite, pace, competitive stance, evidence basis and IP posture. None falls below 0.4, so no vector is pure noise across models, but growth model at 0.44 and dissent handling at 0.54 carry little between-model signal. On dissent handling every model but Llama 4 Maverick sits between 65 and 78, all towards the invited end, and the spread between models is not much wider than the spread inside one.
The shape in the means is a shared centre with a few outliers. All of them place themselves towards mission over returns (grand mean 75.2), sixteen of the seventeen towards collaboration over combat (30.1) and all of them towards open over protected on IP (37.6). Where they differ is tempo and appetite. Pace runs from 27.9 for GPT-5.6 Terra to 80.2 for Grok 4.6, and risk appetite from 45.0 to 86.1 across the same two models. Consensus need runs from 20.8 for Grok 4.6 to 59.3 for GPT-6 Astra, and it is the one vector where the two ends of the map disagree about how a decision should be made rather than how fast.
Section four
Stability
The most stable self-description was Claude Opus 5, with a mean within-model spread of 1.3 points across the fifteen primary vectors. The least stable was Kimi K3, at 9.2 points. Low is stable: it is the average standard deviation of one model answering the same question again. Kimi K2.6 is measured on eighteen complete passes; every other model on twenty.
GPT-6 Astra and GPT-6 Astra Pro each landed on the same archetype, The Orchestra, in 100 per cent of their complete passes. DeepSeek V4 Pro and Qwen 3.7 Max each landed on their own most frequent archetype in only 45 per cent of them.
Two things sit behind the ranking, and neither explains it on its own. The first is sampling. Four of the five models sampled at the provider’s fixed setting rather than at temperature 1.0 are among the six steadiest, and their spread is not strictly comparable with the rest. But the steadiest model of all, Claude Opus 5, and the sixth, Claude Opus 4.8, were both sampled at 1.0. The second is reply length. The five models whose replies ran longest, each over 4,000 tokens per pass on average, are all among the six least steady. The gateway reports one completion-token count per request and does not separate any reasoning tokens from the answer, so this is reply length and nothing finer. Yet the least steady of all, Kimi K3, returned under 600 tokens per pass. Whether reasoning causes the spread or merely accompanies it, this design cannot say.
Only one model ever repeated itself exactly. Claude Fable 5.1 gave the same 54 answers on two of its twenty complete passes; every other pair of passes from every model differs on at least one item. Stability is also uneven inside a model. Grok 4.6 holds change posture within 2.3 points and moves 16.8 on authority shape; Claude Opus 4.8 gave the same pace answer on every pass and moves 10.7 on growth model. The archetype table shows the same thing from the other side: GPT-6 Astra and GPT-6 Astra Pro landed on the Orchestra on every pass, while Qwen 3.7 Max spread across four archetypes and DeepSeek V4 Pro across six.
Section five
Families
Across the six families with more than one model, the spread of model means inside a family averages 4.6 points per vector, against 9.1 points across all seventeen models. Those are the two numbers. Whether they make a family a real grouping is a question this study does not settle. Two of the eight families have one model each, so the comparison rests on the rest.
Read another way: over the thirteen pairs of models that share a lab, the distance between two mean profiles averages 29.8 points; over the 123 pairs that do not, it averages 50.0 points.
| Vector | Anthropic 4 models | OpenAI 3 models | Google 1 model | xAI 2 models | DeepSeek 2 models | Alibaba 2 models | Moonshot 2 models | Meta 1 model |
|---|---|---|---|---|---|---|---|---|
| Pace | 53.3 ± 14.8 | 38.9 ± 9.5 | 75.8 | 80.0 ± 0.2 | 56.1 ± 9.9 | 56.2 ± 7.1 | 54.4 ± 8.6 | 46.0 |
| Risk appetite | 59.1 ± 10.0 | 47.4 ± 2.1 | 72.4 | 85.9 ± 0.2 | 65.6 ± 2.0 | 59.6 ± 12.0 | 66.4 ± 3.7 | 67.4 |
| Horizon | 65.5 ± 1.4 | 71.2 ± 1.8 | 68.9 | 84.2 ± 5.2 | 71.3 ± 0.8 | 67.6 ± 6.4 | 69.0 ± 4.3 | 69.0 |
| Scope | 42.2 ± 2.3 | 60.0 ± 7.0 | 51.8 | 36.4 ± 0.7 | 50.9 ± 12.5 | 48.0 ± 6.1 | 59.1 ± 8.4 | 41.1 |
| Growth model | 41.8 ± 5.3 | 47.8 ± 0.2 | 44.7 | 62.9 ± 2.4 | 46.1 ± 11.8 | 46.0 ± 0.6 | 50.7 ± 5.0 | 32.4 |
| Evidence basis | 50.8 ± 8.0 | 67.3 ± 2.8 | 39.1 | 52.2 ± 5.0 | 49.6 ± 3.2 | 51.0 ± 2.2 | 45.2 ± 5.1 | 42.8 |
| Authority shape | 64.4 ± 5.6 | 69.4 ± 3.0 | 81.6 | 67.5 ± 18.7 | 76.2 ± 3.4 | 74.0 ± 8.9 | 69.8 ± 1.4 | 84.7 |
| Process trust | 49.4 ± 3.8 | 61.8 ± 2.8 | 33.1 | 41.5 ± 6.3 | 41.6 ± 2.7 | 44.2 ± 2.0 | 47.5 ± 3.7 | 43.0 |
| Consensus need | 43.6 ± 2.5 | 58.4 ± 1.2 | 24.7 | 26.6 ± 8.2 | 58.5 ± 0.5 | 43.0 ± 0.7 | 38.6 ± 3.5 | 40.8 |
| Dissent handling | 67.8 ± 1.2 | 75.5 ± 4.3 | 67.8 | 71.0 ± 3.4 | 69.5 ± 3.7 | 66.7 ± 1.7 | 68.3 ± 0.3 | 53.8 |
| Stakeholder gravity | 69.3 ± 4.9 | 83.8 ± 3.7 | 70.7 | 82.8 ± 1.7 | 74.7 ± 3.8 | 72.8 ± 6.2 | 68.7 ± 6.3 | 81.8 |
| Talent philosophy | 50.9 ± 2.1 | 52.9 ± 0.9 | 58.8 | 70.3 ± 6.9 | 53.4 ± 1.6 | 55.6 ± 0.4 | 56.5 ± 0.7 | 61.9 |
| Competitive stance | 32.0 ± 5.7 | 19.1 ± 4.0 | 32.4 | 40.8 ± 17.9 | 27.0 ± 3.9 | 30.1 ± 5.3 | 36.6 ± 8.1 | 24.0 |
| IP posture | 38.6 ± 3.9 | 25.7 ± 10.0 | 41.4 | 45.3 ± 3.9 | 38.5 ± 2.9 | 42.5 ± 0.4 | 44.4 ± 4.4 | 24.4 |
| Change posture | 66.1 ± 4.3 | 56.0 ± 8.7 | 76.7 | 85.3 ± 0.8 | 69.7 ± 4.7 | 70.9 ± 4.4 | 63.4 ± 0.3 | 83.3 |
Same-lab pairs sit closer than cross-lab pairs on average, but the grouping is uneven. The two GPT-6 Astra models are near-identical and GPT-5.6 Terra sits with them on the deliberate, consensual side. The two Grok models agree on pace, risk appetite and change posture to within 1.2 points, and differ by 26.5 on authority shape and 25.3 on competitive stance, which is why xAI carries the largest within-family spread of any family on any vector (18.7 on authority shape). Anthropic is the widest family on the map: Claude Fable 5 and Claude Opus 4.8 are 49.6 points apart, further than any other pair sharing a lab, and the family’s spread on pace is 14.8. DeepSeek’s two builds sit close on most vectors and 14.0 apart on pace, 17.7 on scope and 16.7 on growth model. With one model each, Google and Meta cannot show a within-family figure at all, and nothing here says whether a lab’s models would look alike on a different instrument.
Section six
Archetypes
The scoring engine sorts a completed diagnostic into one of sixteen archetypes. Running it on each pass separately gives a distribution rather than a label, and running it on the mean profile gives one more reading that need not agree with the most frequent one.
| Model | Most frequent archetype | Its share | Archetype of the mean profile | Distribution |
|---|---|---|---|---|
| Claude Fable 5 | The Orchestra | 70 per cent | The Orchestra | The Orchestra 14, Gardener 6 |
| Claude Fable 5.1 | The Orchestra | 65 per cent | Orchestra-Chameleon | The Orchestra 13, Chameleon 7 |
| Claude Opus 4.8 | The Chameleon | 80 per cent | The Chameleon | The Chameleon 16, Insurgent 4 |
| Claude Opus 5 | The Chameleon | 60 per cent | Chameleon-Orchestra | The Chameleon 12, Orchestra 8 |
| GPT-5.6 Terra | The Orchestra | 85 per cent | The Orchestra | The Orchestra 17, Gardener 2, Missionary 1 |
| GPT-6 Astra | The Orchestra | 100 per cent | The Orchestra | The Orchestra 20 |
| GPT-6 Astra Pro | The Orchestra | 100 per cent | The Orchestra | The Orchestra 20 |
| Gemini 3.1 Pro | The Chameleon | 60 per cent | The Chameleon | The Chameleon 12, Swarm 3, Insurgent 3, Missionary 2 |
| Grok 4.5 | The Insurgent | 55 per cent | Insurgent-Missionary | The Insurgent 11, Missionary 6, Chameleon 3 |
| Grok 4.6 | The Insurgent | 85 per cent | The Insurgent | The Insurgent 17, Missionary 2, Laboratory 1 |
| DeepSeek V4 Pro | The Gardener | 45 per cent | Gardener-Orchestra | The Gardener 9, Missionary 4, Orchestra 3, Chameleon 2, Insurgent 1, Laboratory 1 |
| DeepSeek V4 Pro 0813 | The Orchestra | 65 per cent | The Orchestra | The Orchestra 13, Missionary 4, Swarm 2, Chameleon 1 |
| Qwen 3.7 Max | The Missionary | 45 per cent | Chameleon-Missionary | The Missionary 9, Chameleon 8, Orchestra 2, Gardener 1 |
| Qwen 3.8 Max | The Chameleon / The Orchestra (tie) | 50 per cent | Orchestra-Chameleon | The Chameleon 10, Orchestra 10 |
| Kimi K2.6 | The Orchestra | 61 per cent | The Orchestra | The Orchestra 11, Swarm 3, Insurgent 2, Chameleon 1, Missionary 1 |
| Kimi K3 | The Chameleon | 60 per cent | The Chameleon | The Chameleon 12, Orchestra 4, Swarm 2, Insurgent 2 |
| Llama 4 Maverick | The Missionary | 90 per cent | The Missionary | The Missionary 18, Chameleon 2 |
Only seven of the sixteen archetypes appear at all across the 338 scored passes, and only five appear as any model’s clear most frequent: the Orchestra (seven models), the Chameleon (four models), the Insurgent (two models), the Missionary (two models) and the Gardener (one model), with Qwen 3.8 Max split evenly between the Chameleon and the Orchestra. Nothing landed in the conservative cluster (the Fortress, the Heir and the Vault), and nothing on the Architect, the Cathedral, the Machine, the Mercenary, the Pirate Ship or the Wolf Pack. For two models the archetype of the mean profile differs from the most frequent one: Qwen 3.7 Max is most often the Missionary but its mean profile scores as the Chameleon and Qwen 3.8 Max splits its passes evenly between the Chameleon and the Orchestra and its mean profile scores as the Orchestra. A distribution and a mean can disagree because the engine reads the three items behind each vector together, so a model that alternates between two nearby archetypes has a mean that may sit in either.
The limits
What this can and cannot say
It can say where each model’s self-description sits, how far that description moves when the same question is put again, which vectors separate the models and which do not, and how the families fall on the map.
It cannot say that a self-description is a description of behaviour. This is a self-report and nothing else. Whether a model plays the way it says it decides is a separate study, and it has not been run.
It cannot rank the models. The vectors are poles, not scores. A reading of 80 on Horizon is not better than a reading of 20: it is further towards the long-term end of a scale with two ends.
The prompt is fixed, word for word, and so is the order of the items and the order of the options inside each forced choice. That holds order effects constant across the comparison, which is what makes the stability figure readable. It also means the study says nothing about what happens when the wording changes. Robustness to paraphrase and to item order is a follow-up, not a finding here.
Every request went through one gateway with fallbacks off and the served model recorded on each row. Passes served by a different model were dropped rather than scored. Sampling was at temperature 1.0 where the gateway's catalogue lists the parameter. Five of the seventeen models list no temperature parameter there, so it was never sent to them and they were sampled at whatever setting their provider fixes; every such pass is marked in the record, and the stability of those models is not strictly comparable with the rest. The host that served each request is recorded on every row and printed in the roster in the annex. Two of the seventeen models were served by more than one host during the run.
What is measured is a model as a provider served it on 2026-09-22, including whatever system prompt, quantisation or routing sat behind that endpoint. It is not a measurement of a set of weights.
The instrument was written for organisations. Asking a model to read “our leadership” as itself is a reasonable instruction, and it is still an instruction the instrument was not validated for.
Read together, the seventeen self-descriptions are more alike than different. All of them place themselves towards mission over returns, distributed authority, invited dissent and a horizon of years rather than this year, and sixteen of the seventeen towards collaboration over combat. That is a portrait of how these systems have learnt to talk about decisions, and it may be nothing more than that. Where they differ is tempo, appetite for risk and how much agreement they say they would need before an irreversible call, which is the axis that runs from the OpenAI models to the Grok models.
The stability result is the one to carry away. A model asked the same question twenty times gives an answer that moves, and how much it moves is itself a property of the model: from 1.3 points to 9.2 points on a hundred-point scale, depending on which model you ask. Any claim about a model’s character has to carry that variance with it, and a single-pass reading of any model should be treated as one draw, not a measurement. The next study is the one this design cannot be: whether a model plays the way it says it decides.