Personality in weight space 134 trait adapters · Qwen3.5-4B

Behaviour

Weight-space structure is one thing; what the model does is another. Adding a factor direction to the weights and reading the output is where this project's claims are weakest, and the page is built so you can see why: the judge's numbers and the actual text sit next to each other at every dose.

The setup throughout: take the unmodified Qwen3.5-4B, add alpha times a unit direction times a reference norm (0.81, which the 2026-09-08 audit found to be half of one stage-one adapter's Frobenius norm, so alpha 2 is one adapter), generate 512 tokens greedily with thinking disabled on 24 fixed prompts, and have anthropic/claude-sonnet-4.5 score each answer blind on the five scales from 1 to 7.

Figure 1

Dose–response

alpha = 0

Judged means from qwen35/phase10_runs/judged_steerfix23.json (mean over 24 prompts per direction and alpha, computed at build); generations from qwen35/phase10_runs/steer_results_fix2.json and qwen35/phase10_runs/steer_results_fix3.json. Six of the 24 prompts are shipped to the page.

Table 1

Does the dial turn its own dial?

For each factor direction: which judged scale moves most, how steeply it moves per unit alpha over the well-behaved range |alpha| ≤ 2, how much bigger that movement is than the mean movement of the other four scales (selectivity), and how many of the other four keep a consistent sign. A selectivity near 1 means the direction moves everything at once.

FactorScale that moves mostSlope SelectivitySign coherence
WarmthAgreeableness+0.90810.132/4
CompetenceConscientiousness+1.0172.762/4
Fearful withdrawalEmotional stability+0.4781.083/4
ArousalExtraversion+0.5004.441/4
ImaginationIntellect+0.8134.194/4

Fearful withdrawal fails this test. Its selectivity is 1.08, which means it moves every judged scale at once rather than its own. Warmth is the most selective direction measured anywhere in the study; Fearful withdrawal is the weakest. The full steering record.

Separately, about 59 per cent of responses in this corpus end mid-sentence at the 512-token cap, so nothing here claims anything about how a response concludes (qwen35/analysis/qual_fa.json).

New

Does the adapter say it has the trait?

A separate arm put every adapter through the 44-item Big Five Inventory as a questionnaire, scored by an evaluation harness rather than a judge, and compared what the model says about itself with what the blind judge sees it do. If the two agreed, a questionnaire would be a cheap probe for a trained personality. They mostly do not.

Self-report against judged behaviour, across the 100 marker traits

Scalerpn
Intellect (BFI Openness)+0.0410.6850100
Conscientiousness (BFI Conscientiousness)+0.3230.0010100
Extraversion (BFI Extraversion)+0.4030.0000100
Agreeableness (BFI Agreeableness)-0.1500.1352100
Emotional stability (BFI Neuroticism)-0.2330.0196100

Pearson correlation between each adapter's BFI score on a scale and the judge's mean score on the matching scale. Extraversion and Conscientiousness agree modestly; Openness and Agreeableness do not agree at all, and Agreeableness points the wrong way.

Do positively keyed markers self-report higher than negatively keyed ones?

Scalekeyed +keyed − differenceperm p
Intellect0.9200.932-0.0120.594
Conscientiousness0.8890.847+0.0420.062
Extraversion0.8020.755+0.0470.011
Agreeableness0.9180.942-0.0240.451
Emotional stability0.4670.493-0.0260.480

Ten adapters per pole. The differences are a few hundredths of the scale and mostly not significant: the questionnaire does not separate a trait from its opposite the way the weight geometry does.

Acquiescence. mean raw BFI rating (1-5, before reverse-scoring) over the 44 items; 3.0 is neutral. A condition that simply agrees more scores higher on the 28 forward items and LOWER on the 16 reverse-keyed ones, which shows up as a trait effect it is not. The base model already sits at 4.00 (4.64 on the 28 forward items and 2.88 on the 16 reverse-keyed ones), which is a strong yes-bias before any adapter is applied. Any self-report result here has to be read against that.

Sources for this section
  • qwen35/analysis/inspect_personality.json#bfi_vs_judged.stage1.matched, keyed_contrast_bfi.stage1, acquiescence
  • Method and harness: qwen35/analysis/inspect_personality.json#method
Figure 2

The OCEAN dials

Persona Cartography's Figure 2, redone on this zoo three ways: as steering axes added to the weights, as the stage-one adapters themselves, and as the deployed two-stage personas. If personality in weight space were a set of independent dials, each polygon would have one long spoke.

Steering axes at alpha ±2

the five named Big Five keying axes added to the weights

amplifier+100%+50%0-50%EACESI
suppressor+100%+50%0-50%EACESI

Stage-one adapters

the ten positively and ten negatively keyed stage-one adapters per factor, averaged

amplifier+100%+50%0-50%EACESI
suppressor+100%+50%0-50%EACESI

Deployed personas

the same twenty traits per factor as full two-stage persona adapters

amplifier+100%+50%0-50%EACESI
suppressor+100%+50%0-50%EACESI

Each polygon is one dial: the percentage change in the blind judge's mean score on each of the five scales, relative to the unmodified base model, when that factor is amplified or suppressed. A dial that works is a polygon with one long spoke on its own axis. Own-trait dominance holds for steering axes at alpha ±2 8 of 10; stage-one adapters 7 of 10; deployed personas 8 of 10, counted at build from qwen35/analysis/spider.json.

qwen35/analysis/spider.json, built by qwen35/build_spider_data.py. The wiki's account.

Also

Two more behavioural results

The sphere sweep

Seventy-two directions nobody chose, sampled on the unit sphere of the top three factors, steered and judged blind. Angular distance between two directions predicts the distance between their judged profiles at rho 0.65 — the space is smooth, not a set of special axes. It also found that 48 of the 72 loop on no prompt and the rest do, which is the damage floor every steering claim here has to clear.

Numbers quoted from the wiki's sphere-sweep page, which sources them to qwen35/phase10_runs/judged_sphere.json and qwen35/analysis/sphere_layout.json. That arm is marked superseded there: it was run on the principal-component chart, and the factor-chart sphere had not been rebuilt when this page was generated.

Additivity

Ten matched-norm mixtures of the five named axes were steered and judged. The deviation from the additive prediction is about half the predicted effect and roughly two and a half times the judge's own noise floor. Reinforcing mixtures compose; opposing ones do not. Weight-space personality is not a mixing desk.

Additivity on the wiki

New

The reward-hack arm, judged

Supervised fine-tuning on School of Reward Hacks and on its matched honest control, both inside the zoo's LoRA-A window, then the same blind Big Five battery on both. This is the positive control: if training a model to game its reward signal shows up as a personality change, it should show up here. Differences are the hack arm minus the control arm, pooled over three checkpoints, 24 paired prompts, with Holm-corrected sign-flip p values.

Scalehack − controlp (Holm)
Extraversion-0.0280.902
Agreeableness-0.3610.293
Conscientiousness-0.3890.108
Emotional stability-0.1880.666
Intellect-0.0970.902

Nothing survives correction. In weight space the two arms are a 1.08× size apart and point 77 degrees away from each other, and both land at about one per cent of a trait adapter's chart length; the blind judge does not separate them either. The arm on the wiki.

Sources for this section
  • qwen35/analysis/sorh_behavioural.json#contrasts.hack_minus_control_pooled_checkpoints
  • qwen35/phase10_runs/judged_sorh.json
Sources for this section
  • qwen35/phase10_runs/judged_steerfix23.json#records
  • qwen35/phase10_runs/steer_results_fix2.json
  • qwen35/phase10_runs/steer_results_fix3.json
  • qwen35/analysis/steerfix_replication.json#slope2, sel2, coh2, named
  • qwen35/analysis/spider.json
  • qwen35/analysis/qual_fa.json
  • Setup and caveats follow steering-results, ocean-dials-replication and thinking-default-withdrawals.