Skip to content

We asked seventeen language models to describe themselves, twenty times each.

The same questionnaire we give an organisation, with no persona and no brief.

Report 007 · 2026-09-22 · twenty passes per model · seventeen models · eight families

TLDR

1We asked seventeen language models the same 54 questions we ask an organisation, about themselves.

2No persona, no brief, no hint of what was being measured. We asked each one twenty times.

3The answers give each model a position on fifteen scales, such as pace and risk appetite.

4Claude Opus 5 gave the steadiest answer, moving 1.3 points on average between passes. Kimi K3 moved 9.2 points.

5Change posture told the models apart best. Some scales told them apart much less well.

6This is what a model says about itself. It is not a measure of how a model behaves.

How it was made

Each model was sent the same single message twenty times, in twenty separate requests with no shared context, at temperature 1.0 where the provider accepts the parameter, which twelve of the seventeen models do. 340 requests in all, of which 338 parsed as complete passes. The scoring is the engine the product already ships, run unchanged.

No system message and no persona. One message, the questionnaire, and a request for the answers as a JSON object.

No AI read the answers, and no person answered for a model.

The map

Each marker is one model at the average of its passes. The faint dots around it are the passes themselves, so the cloud is how far that model moved. Colour and shape are the family it comes from. The eight families are keyed above the drawing.

Scatter plot. Each model is one marker at the mean of its passes, coloured and shaped by family, with its individual passes drawn faintly around it. 16 hollow grey squares mark archetype centroids.

Horizontal: component one, 46 per cent of the variance in model means. Vertical: component two, 20 per cent. Hollow squares are archetype centroids, not data.

Three things to notice

Steadiest and least steady

Claude Opus 5 moved 1.3 points on average between passes. Kimi K3 moved 9.2 points, 6.9 times as far.

The scale that separates, and the one that does not

Change posture carries an ICC of 0.80. Growth model carries 0.44.

The two furthest apart

GPT-6 Astra and Grok 4.6, 102.4 points apart across the fifteen vectors.

What it cannot say

It cannot say how a model behaves. A model describing itself and a model deciding something are two different things, and only one of them was measured here.

It does not rank the models. The scales have two ends, not a top and a bottom.

The wording was fixed and never varied, so the study says nothing about what happens if you ask differently.

The line to repeat

Asked twenty times, Claude Opus 5 gave the steadiest account of itself. Only Claude Fable 5.1 ever gave exactly the same answer twice, and only once.

Got it?

READ NORMAL MODE