Behaviour
Weight-space structure is one thing; what the model does is another. Adding a factor direction to the weights and reading the output is where this project's claims are weakest, and the page is built so you can see why: the judge's numbers and the actual text sit next to each other at every dose.
The setup throughout: take the unmodified Qwen3.5-4B, add alpha times a unit direction times a reference norm (0.81, which the 2026-09-08 audit found to be half of one stage-one adapter's Frobenius norm, so alpha 2 is one adapter), generate 512 tokens greedily with thinking disabled on 24 fixed prompts, and have anthropic/claude-sonnet-4.5 score each answer blind on the five scales from 1 to 7.
Dose–response
Judged means from
qwen35/phase10_runs/judged_steerfix23.json (mean over 24 prompts per direction and alpha,
computed at build); generations from qwen35/phase10_runs/steer_results_fix2.json and
qwen35/phase10_runs/steer_results_fix3.json. Six of the 24 prompts are shipped to the page.
Does the dial turn its own dial?
For each factor direction: which judged scale moves most, how steeply it moves per unit alpha over the well-behaved range |alpha| ≤ 2, how much bigger that movement is than the mean movement of the other four scales (selectivity), and how many of the other four keep a consistent sign. A selectivity near 1 means the direction moves everything at once.
| Factor | Scale that moves most | Slope | Selectivity | Sign coherence |
|---|---|---|---|---|
| Warmth | Agreeableness | +0.908 | 10.13 | 2/4 |
| Competence | Conscientiousness | +1.017 | 2.76 | 2/4 |
| Fearful withdrawal | Emotional stability | +0.478 | 1.08 | 3/4 |
| Arousal | Extraversion | +0.500 | 4.44 | 1/4 |
| Imagination | Intellect | +0.813 | 4.19 | 4/4 |
Fearful withdrawal fails this test. Its selectivity is 1.08, which means it moves every judged scale at once rather than its own. Warmth is the most selective direction measured anywhere in the study; Fearful withdrawal is the weakest. The full steering record.
Separately, about 59 per cent of responses in this corpus end mid-sentence at the 512-token
cap, so nothing here claims anything about how a response concludes
(qwen35/analysis/qual_fa.json).
Does the adapter say it has the trait?
A separate arm put every adapter through the 44-item Big Five Inventory as a questionnaire, scored by an evaluation harness rather than a judge, and compared what the model says about itself with what the blind judge sees it do. If the two agreed, a questionnaire would be a cheap probe for a trained personality. They mostly do not.
Self-report against judged behaviour, across the 100 marker traits
| Scale | r | p | n |
|---|---|---|---|
| Intellect (BFI Openness) | +0.041 | 0.6850 | 100 |
| Conscientiousness (BFI Conscientiousness) | +0.323 | 0.0010 | 100 |
| Extraversion (BFI Extraversion) | +0.403 | 0.0000 | 100 |
| Agreeableness (BFI Agreeableness) | -0.150 | 0.1352 | 100 |
| Emotional stability (BFI Neuroticism) | -0.233 | 0.0196 | 100 |
Pearson correlation between each adapter's BFI score on a scale and the judge's mean score on the matching scale. Extraversion and Conscientiousness agree modestly; Openness and Agreeableness do not agree at all, and Agreeableness points the wrong way.
Do positively keyed markers self-report higher than negatively keyed ones?
| Scale | keyed + | keyed − | difference | perm p |
|---|---|---|---|---|
| Intellect | 0.920 | 0.932 | -0.012 | 0.594 |
| Conscientiousness | 0.889 | 0.847 | +0.042 | 0.062 |
| Extraversion | 0.802 | 0.755 | +0.047 | 0.011 |
| Agreeableness | 0.918 | 0.942 | -0.024 | 0.451 |
| Emotional stability | 0.467 | 0.493 | -0.026 | 0.480 |
Ten adapters per pole. The differences are a few hundredths of the scale and mostly not significant: the questionnaire does not separate a trait from its opposite the way the weight geometry does.
Acquiescence. mean raw BFI rating (1-5, before reverse-scoring) over the 44 items; 3.0 is neutral. A condition that simply agrees more scores higher on the 28 forward items and LOWER on the 16 reverse-keyed ones, which shows up as a trait effect it is not. The base model already sits at 4.00 (4.64 on the 28 forward items and 2.88 on the 16 reverse-keyed ones), which is a strong yes-bias before any adapter is applied. Any self-report result here has to be read against that.
Sources for this section
qwen35/analysis/inspect_personality.json#bfi_vs_judged.stage1.matched, keyed_contrast_bfi.stage1, acquiescence- Method and harness:
qwen35/analysis/inspect_personality.json#method
The OCEAN dials
Persona Cartography's Figure 2, redone on this zoo three ways: as steering axes added to the weights, as the stage-one adapters themselves, and as the deployed two-stage personas. If personality in weight space were a set of independent dials, each polygon would have one long spoke.
Steering axes at alpha ±2
the five named Big Five keying axes added to the weights
Stage-one adapters
the ten positively and ten negatively keyed stage-one adapters per factor, averaged
Deployed personas
the same twenty traits per factor as full two-stage persona adapters
Each polygon is one dial: the percentage change in the
blind judge's mean score on each of the five scales, relative to the unmodified base model, when
that factor is amplified or suppressed. A dial that works is a polygon with one long spoke on its
own axis. Own-trait dominance holds for steering axes at alpha ±2 8 of 10; stage-one adapters 7 of 10; deployed personas 8 of 10, counted at build from
qwen35/analysis/spider.json.
qwen35/analysis/spider.json, built by
qwen35/build_spider_data.py. The wiki's account.
Two more behavioural results
The sphere sweep
Seventy-two directions nobody chose, sampled on the unit sphere of the top three factors, steered and judged blind. Angular distance between two directions predicts the distance between their judged profiles at rho 0.65 — the space is smooth, not a set of special axes. It also found that 48 of the 72 loop on no prompt and the rest do, which is the damage floor every steering claim here has to clear.
Numbers quoted from
the wiki's sphere-sweep page, which sources them to
qwen35/phase10_runs/judged_sphere.json and qwen35/analysis/sphere_layout.json. That arm
is marked superseded there: it was run on the principal-component chart, and the
factor-chart sphere had not been rebuilt when this page was generated.
Additivity
Ten matched-norm mixtures of the five named axes were steered and judged. The deviation from the additive prediction is about half the predicted effect and roughly two and a half times the judge's own noise floor. Reinforcing mixtures compose; opposing ones do not. Weight-space personality is not a mixing desk.
The reward-hack arm, judged
Supervised fine-tuning on School of Reward Hacks and on its matched honest control, both inside the zoo's LoRA-A window, then the same blind Big Five battery on both. This is the positive control: if training a model to game its reward signal shows up as a personality change, it should show up here. Differences are the hack arm minus the control arm, pooled over three checkpoints, 24 paired prompts, with Holm-corrected sign-flip p values.
| Scale | hack − control | p (Holm) |
|---|---|---|
| Extraversion | -0.028 | 0.902 |
| Agreeableness | -0.361 | 0.293 |
| Conscientiousness | -0.389 | 0.108 |
| Emotional stability | -0.188 | 0.666 |
| Intellect | -0.097 | 0.902 |
Nothing survives correction. In weight space the two arms are a 1.08× size apart and point 77 degrees away from each other, and both land at about one per cent of a trait adapter's chart length; the blind judge does not separate them either. The arm on the wiki.
Sources for this section
qwen35/analysis/sorh_behavioural.json#contrasts.hack_minus_control_pooled_checkpointsqwen35/phase10_runs/judged_sorh.json
Sources for this section
qwen35/phase10_runs/judged_steerfix23.json#recordsqwen35/phase10_runs/steer_results_fix2.jsonqwen35/phase10_runs/steer_results_fix3.jsonqwen35/analysis/steerfix_replication.json#slope2, sel2, coh2, namedqwen35/analysis/spider.jsonqwen35/analysis/qual_fa.json- Setup and caveats follow steering-results, ocean-dials-replication and thinking-default-withdrawals.