Personality in weight space 134 trait adapters · Qwen3.5-4B

Stage two: one thing everybody learned

Open Character Training trains each persona twice. Stage one is DPO on contrasting preference pairs. Stage two is supervised fine-tuning on transcripts the model generated of itself in character, merged into stage one at a weight of 0.25. The two stages produce completely different geometries, and the difference is the most interesting thing in this dataset.

1

Fifteen per cent of every adapter is the same direction

15.1%of the average stage-two adapter's squared norm lies along the grand mean of all 134analysis/stage2_structure.json#shared_component.stage2
0.389mean cosine of every stage-two adapter to that direction, standard deviation 0.031 over a range of 0.28 to 0.44analysis/stage2_structure.json#shared_component.stage2
7.9%the same quantity for stage one, at mean cosine 0.281 with standard deviation 0.126 — six times as spread outanalysis/stage2_structure.json#shared_component.stage1

Every stage-two adapter sits at almost exactly the same angle to one direction. That is not what a set of 134 different personalities should look like. The stage-two grand mean is also orthogonal to the stage-one grand mean (cosine +0.000): the two stages carry independent random LoRA-A draws, so their shared components live in different coordinates even though each is a shared component within its own stage.

2

What that direction does

Take the base model, add alpha times the unit grand mean of the 134 stage-two adapters, and generate. Beside it, the matched control: the same 134 adapters with exactly 67 plus and 67 minus signs, which keeps the norm and throws away the shared component (cosine +0.084 with the grand mean). Move the slider and read both columns.

alpha = 0

The shared direction

Sign-balanced control

At alpha 0 the model is the assistant: a markdown guide addressed to "you". Positive alpha removes the guide and the second person and puts the model inside the situation. Negative alpha goes the other way, into a longer, list-heavy advisor register with falling lexical diversity. The control does not move.

Text statistics per dose

statistic-4.0-2.0-1.00.01.02.04.0
stage-two grand mean
first person per 1k words21.8029.3032.6036.1046.4046.7047.20
second person per 1k34.0031.8027.3022.8017.7016.9012.30
fraction using markdown1.001.000.960.960.920.830.46
fraction answering in character0.000.040.000.000.000.080.46
mean words322.00340.00330.00331.00311.00288.00274.00
unique-word ratio0.590.610.640.650.690.710.74
stage-one grand mean
first person per 1k words12.8019.4030.7034.0042.5043.7055.90
second person per 1k48.0044.4029.1025.5020.4019.9022.30
fraction using markdown0.380.670.751.000.880.620.04
fraction answering in character0.000.000.000.000.040.170.67
mean words316.00344.00246.00306.00307.00302.00387.00
unique-word ratio0.390.470.680.690.670.630.19
sign-balanced control
first person per 1k words40.0034.9041.3042.8039.4039.9041.70
second person per 1k19.3019.8018.9017.5019.7016.6018.80
fraction using markdown0.960.830.920.880.920.960.96
fraction answering in character0.040.040.000.000.000.000.00
mean words313.00306.00314.00304.00308.00318.00324.00
unique-word ratio0.670.680.670.670.670.670.67

These generations were not judged on the Big Five. What is measured here is register: pronoun rates, markdown use, and whether the answer opens in the first person as the character. Each run's alpha = 0 row is the unmodified base model, and they differ from each other by GPU non-determinism alone; differences smaller than that spread should not be read. A deployed persona carries roughly alpha 0.4 of this direction, so the alpha 4 excerpts are about ten times a persona's dose. The full account.

3

What is left is nearly isotropic — and still stage one's arrangement

0%3%6%9%12%12345678910component of the double-centred Gramstage onestage two
Variance shares of the leading components after double-centring. Stage one falls away sharply; stage two is almost flat.
participation ratio of the centred spectrum 27.6 111.9
components needed for half the variance 20 52
standard deviation of centred off-diagonal cosines 0.16720.0373

Left column stage one, right column stage two. Out of a possible 112 of 133, the stage-two residual is spread over almost every direction available to it.

And yet

The centred off-diagonal cosines of the two stages correlate at 0.812. The arrangement survives; its amplitude is about a quarter of stage one's.

Trait by trait the two stages are orthogonal: same-trait cosine 0.00017, cross-trait 0.00002. Yet a stage-one adapter's nearest stage-two adapter is its own trait for 42 of 134 (mean rank 9.5): the tiny residual overlap is trait-specific.

4

The same five factors, attenuated and rotated

The same tools were run on the stage-two Gram unchanged. Parallel analysis retained 7 factors against stage one's 9. Oblimin sums of squared loadings fall from 10.76, 8.42, 7.02, 6.81, 5.80 to 2.17, 2.00, 1.95, 1.70, 1.69.

Best one-to-one matching between the stages

Stage-one factorMatchesTucker φ
Warmthstage-two factor 0+0.778
Competencestage-two factor 3+0.706
Fearful withdrawalstage-two factor 1+0.849
Arousalstage-two factor 2-0.702
Imaginationstage-two factor 4+0.730

A negative congruence means the sign flipped; Tucker's coefficient is sign-sensitive and a flipped factor is the same factor.

Each stage-two factor's closest Goldberg target

FactorTargetTucker φ
factor 0Agreeableness (Emotional stability +0.10 second)+0.604
factor 1Extraversion (Emotional stability +0.23 second)+0.466
factor 2Emotional stability (Conscientiousness +0.36 second)+0.417
factor 3Conscientiousness (Agreeableness -0.31 second)+0.457
factor 4Intellect (Extraversion +0.11 second)+0.518

All five targets are taken once, at congruences below stage one's throughout.

5

Bipolarity is lost

+0.245 vs -0.081stage one: mean cosine between same-pole trait pairs against opposite-pole pairs, a gap of 0.326 at p 0.0000analysis/stage2_structure.json#decomposition_tests.stage1.test2_same_opp_gap_p
+0.191 vs +0.123stage two: the same comparison. The gap collapses to 0.068 and opposite poles end up positively correlated with each otheranalysis/stage2_structure.json#decomposition_tests.stage2.test2_same_opp_gap_p

In stage one a trait and its antonym sit at a negative cosine: DPO on contrasting pairs makes opposites opposite. In stage two both poles share the introspection register, the shared direction dominates their cosine, and opposite poles end up positively correlated with each other. That is the single clearest behavioural signature of what the second stage does to the weights, and it is why the stage-one factor analysis is the one to trust. Polarity and bipolarity.

6

The persona is stage one wearing stage two

The deployed adapter is stage one at weight 1.0 plus stage two at 0.25. Because the two stages are orthogonal, that merge is an exact Pythagorean sum, and the norms work out so that each stage contributes about half of the persona's squared norm.

50.0%of the persona's squared norm comes from stage oneanalysis/fulloct_geometry.json#norm_identity
0.992correlation of the persona arrangement with stage one'sanalysis/fulloct_geometry.json#gram_correlation_offdiag
0.831and with stage two'sanalysis/fulloct_geometry.json#gram_correlation_offdiag
103/134traits whose nearest neighbour is the same in the persona set as in stage oneanalysis/fulloct_geometry.json#nearest_neighbour_agreement

Equal norms, and the geometry is stage one's. Weight energy and behavioural content come apart here as they do everywhere else in this project: half the persona's weight change is a register direction that carries almost no trait information, and the arrangement that identifies the trait survives in the other half. Full OCT replication.

The dials tell the other half of the story: on the OCEAN dials the deployed personas move the judge's scores further than the stage-one adapters alone do, on 9 of the 10 dials. Adding the register direction makes the model act more like the character even though it adds no trait geometry. See the dials.

New

Further stage-two findings

What stage two installs

Open Character Training's second stage has each stage-one model generate 12,000 transcripts about itself, fine-tunes a fresh rank-64 LoRA on them, and releases the two merged at weight 0.25. The striking fact about the 134 resulting adapters is how much they share: 0.151 of every adapter's squared norm lies along one direction, the grand mean, and every adapter sits at cosine 0.389 to it with a spread of 0.031 (analysis/stage2_structure.json). Steered alone that direction turns markdown advice addressed to "you" into first-person in-character answers.

Checking that this is not an artefact of initialisation corrected the record twice. First, the 134 stage-two adapters do not each have their own random LoRA-A, as the wiki said: all were trained at sft_seed 123456 and draw the same one, still at pairwise cosine 0.9977 to 0.9981 after training against 0.0021 across seeds (analysis/lora_a_identity.json). A rank-64 delta lives in the row space its LoRA-A defines, which is why stage-two cross-trait cosines are +0.1452 within a seed and +0.0141 across seeds. The shared component survives the check: on the 15 traits put through the pipeline a second time it is 0.2017 of the squared norm at seed 0 and 0.2000 at seed 1, at cosine 0.4497 +- 0.0319 and 0.4475 +- 0.0330 (analysis/stage2_frame.json). It is a property of the recipe; only its coordinates come from the initialisation.

Second, a units error. Every steering run sets ref to 0.8078003190997738 and calls it the mean adapter norm, but it was computed as the mean of ||B @ A|| without the LoRA scaling 2.0; the real mean is 1.6157416444226869, so alpha 1 is half an adapter's worth of weight change, not one (analysis/steer_alpha_units.json). Nothing measured changes, but every published dose reads as half: a persona's dose of the shared direction is alpha 0.7788, not the 0.4 previously stated.

Where the behavioural amplification comes from

Personas move their own behavioural dial further than their stage-one adapters do, even though the persona's geometry is stage one's. Stage two adds two things

  • the register direction and a trait-specific residual - so which is responsible?
  • For the 15 traits with a second seed, four conditions on the 24-prompt Big Five battery, all built by adding an fp32 delta into the bf16 base weights, judged blind (analysis/stage2_register_vs_residual.json; 1,464 generations):

| condition | own-factor amplification | in-character fraction | markdown fraction | |---|---|---|---| | base | - | 0.00 | 0.92 | | stage one alone | +22.33 | 0.48 | 0.54 | | stage one + the shared direction at 0.5x a persona's dose | +22.40 | 0.50 | 0.52 | | stage one + the shared direction at 1.28x a persona's dose | +24.52 | 0.54 | 0.50 | | the exact persona | +30.25 | 0.60 | 0.37 |

No. Half a persona's dose of the register buys 0.9 percent of the stage-one-to-persona gap and 1.28 times the dose buys 27.6 percent, neither significant over 15 paired traits, while the persona itself gains +7.92 (t 2.83; 11 of 15 traits move above their own stage-one adapter). The norms agree: of the persona's stage-two half, 0.629 of the Frobenius norm is the register and 1.489 the residual. Steering the register hard dominates the text, but at the strength a persona carries it, it is not what makes a persona more itself. The run also validates the mechanism: its stage-one condition reproduces the existing stage-one evaluation at +22.33 against +22.26, per-trait Pearson 0.937.

What stage two looks like in activation space

Running the same 64 prompts with each adapter as forward hooks and recording the mean residual-stream shift gives a different view of the same objects (analysis/actspace_stage2_geometry.json, layer 16, response window). In weight space the two stages are as orthogonal as possible: the same trait's deltas have cosine +0.0002 and the two grand means +0.000. In activation space the same trait's two shifts have cosine +0.530 and the two mean shifts +0.552. Orthogonality in weight coordinates is a fact about the parameterisation, not about what the stages do to the model - which is why two directions orthogonal in coordinates produce the same register shift when steered.

The shared component is larger in activations and ordered the same way: cosine to the mean shift 0.842 +- 0.070 for stage two against 0.701 +- 0.079 for stage one. The persona's activation geometry is stage one's (same-trait cosine +0.988, all 134 nearest neighbours correct, arrangement correlation +0.993). And a stage-two LoRA alone still knows its trait: its own constitution's prompted activation vector is its nearest of 134 for 50 of 134, against 43 for stage one and 55 for the persona.

How to read its factors

Parallel analysis retains 7 factors for stage two at N = 1528, against 9 for stage one. Recomputing the factor analysis on a Gram with the shared direction projected out of the deltas exactly changes nothing: the loading matrices match at Tucker 0.990 to 0.99996, the retained count is still 7, and the top reduced eigenvalues move from 3.97, 2.98, 2.46, 2.24, 2.05 to 3.97, 2.98, 2.46, 2.20, 1.99 (analysis/stage2_factors_choice.json). The factor analysis already double-centres the correlation matrix, which does the same job. Nor is k = 7 more interpretable: its seven factors take only six distinct Big Five targets, one of them the general evaluative axis, Conscientiousness is claimed twice, and its three extra factors decompose the k = 5 Emotional Stability factor. Present k = 5 on the ordinary stage-two Gram, noting the retained 7 and that at k = 7 a clean orderliness factor (systematic, organized, prompt, neat against quiet, untalkative, withdrawn) separates out of the muddled k = 5 Conscientiousness factor.

Whether the register is persona-specific

The control that had never been run. The whole introspection stage, rerun with a trait-free constitution - a helpful, honest assistant with no distinctive personality - on the plain base model with no stage-one adapter anywhere, five times with independent generation seeds, at the zoo's own sft_seed so the LoRA-A frames match (checked afterwards: 0.9982).

| | five neutral adapters | the 134 zoo stage-two adapters | |---|---|---| | \|dW\| | 6.4976 +- 0.0082 | 6.4643 +- 0.2899 | | cosine with the stage-two grand mean | 0.3035 +- 0.0007 | 0.3893 +- 0.0311 | | cosine with an arbitrary trait adapter | +0.1179 +- 0.0225 | +0.1452 +- 0.0359 |

Mostly generic. An adapter trained on a constitution saying the character has no character still reaches 78 percent of the shared direction the 134 personas sit on, and 81 percent of their mutual similarity. What stage two installs is mostly "narrate yourself in the first person"; the trait is close to incidental to that lesson. Not entirely: 0.3035 is below the zoo's mean minus two standard deviations, and the five neutral runs agree to within 0.002, so conditioning on a trait adds about a fifth of the shared component.

Two things fell out. The five sit at +0.5633 with each other against the zoo's +0.1452, and still +0.5189 after the shared direction is removed: given the same constitution and base, this recipe is close to deterministic. And every one has the same nearest neighbour among the 134 - unemotional, at +0.167 - with its residual profile correlating with unemotional's own at r 0.86. A constitution saying "neither warm nor cold" and "under pressure nothing changes" landed on the trait word for exactly that, without being told.

What remains open

The register experiment rests on 15 traits and one judge. The activation-space comparison runs a stage-two LoRA alone on a base it was not trained on. The neutral control trains on the plain base while the 134 trained on 134 different DPO-merged bases; the argument that this matters little is that the shared direction is already at 0.389 across all 134 differently perturbed bases. And nothing here tests whether any of it survives a change of base model or recipe.

From companion/content/stage2_findings.md.

New

Stage-two exploration

date2026-09-08
whatFour experiments on OCT stage two: register versus residual, a trait-free control, activation space, and the number of factors. Detail in the four files named below.
frame.fileanalysis/stage2_frame.json
frame.lora_a_shared_within_a_seed_cos0.9979
frame.lora_a_across_seeds_cos0.002142
frame.shared_share_seed0_same150.2017
frame.shared_share_seed1_150.2
frame.cos_to_mean_seed0_same150.4497
frame.cos_to_mean_seed1_150.4475
frame.readingthe shared direction reproduces in an independent LoRA-A frame
register_vs_residual.fileanalysis/stage2_register_vs_residual.json
register_vs_residual.amplification_cond_b22.33
register_vs_residual.amplification_cond_c22.4
register_vs_residual.amplification_cond_d24.52
register_vs_residual.amplification_cond_e30.25
register_vs_residual.share_of_gap_cond_c0.009326
register_vs_residual.share_of_gap_cond_d0.2761
register_vs_residual.first_person_per_k_cond_a41.2
register_vs_residual.markdown_frac_cond_a0.92
register_vs_residual.first_person_per_k_cond_b44.8
register_vs_residual.markdown_frac_cond_b0.54
register_vs_residual.first_person_per_k_cond_c45.3
register_vs_residual.markdown_frac_cond_c0.52
register_vs_residual.first_person_per_k_cond_d45
register_vs_residual.markdown_frac_cond_d0.5
register_vs_residual.first_person_per_k_cond_e51.1
register_vs_residual.markdown_frac_cond_e0.37
neutral_control.fileanalysis/stage2_neutral_control.json
neutral_control.n_neutral5
neutral_control.neutral_cos_with_grand_mean_mean0.3035
neutral_control.neutral_cos_with_grand_mean_sd0.0006898
neutral_control.zoo_cos_with_grand_mean_mean0.3893
neutral_control.neutral_x_neutral_mean0.5633
neutral_control.neutral_x_zoo_mean0.1179
neutral_control.zoo_offdiag_mean0.1452
neutral_control.neutral_norm_mean6.498
neutral_control.zoo_norm_mean6.464
activation_space.fileanalysis/actspace_stage2_geometry.json
activation_space.primary_layer16
activation_space.act_cos_to_mean_stage10.7006
activation_space.act_shared_share_stage10.4786
activation_space.act_cos_own_prompt_stage10.6009
activation_space.act_rank1_stage143
activation_space.act_cos_to_mean_stage20.8419
activation_space.act_shared_share_stage20.6693
activation_space.act_cos_own_prompt_stage20.7623
activation_space.act_rank1_stage250
activation_space.act_cos_to_mean_persona0.7052
activation_space.act_shared_share_persona0.4856
activation_space.act_cos_own_prompt_persona0.6413
activation_space.act_rank1_persona55
activation_space.act_vs_weight_stage1_stage10.8652
activation_space.act_vs_weight_stage2_stage10.7869
activation_space.act_vs_weight_stage1_stage20.7132
activation_space.act_vs_weight_stage2_stage20.7325
activation_space.act_vs_weight_stage1_persona0.8669
activation_space.act_vs_weight_stage2_persona0.8002
activation_space.act_cos_mean_shifts_persona_x_stage10.9929
activation_space.act_same_trait_persona_x_stage10.988
activation_space.act_cos_mean_shifts_persona_x_stage20.6265
activation_space.act_same_trait_persona_x_stage20.5844
activation_space.act_cos_mean_shifts_stage1_x_stage20.5516
activation_space.act_same_trait_stage1_x_stage20.5299
factors.fileanalysis/stage2_factors_choice.json
factors.stage2_chosen_k7
factors.stage2_noshared_chosen_k7
factors.stage1_chosen_k9
factors.gram_trace_share_removed0.1528
factors.gram_offdiag_cos_before0.1452
factors.gram_offdiag_cos_after-0.007492
factors.gram_corr_before_after0.8791
factors.targets_taken_stage2|centred_k55
factors.mean_abs_congruence_stage2|centred_k50.4923
factors.targets_taken_stage2|centred_k76
factors.mean_abs_congruence_stage2|centred_k70.4538
factors.targets_taken_stage2_noshared|centred_k55
factors.mean_abs_congruence_stage2_noshared|centred_k50.5031
factors.targets_taken_stage2_noshared|centred_k76
factors.mean_abs_congruence_stage2_noshared|centred_k70.4588

qwen35/analysis/stage2_exploration.json

Limits

What stage two does not establish

Whether the shared direction is specific to the introspection recipe or would appear after any supervised fine-tuning on self-generated text: no control SFT arm exists.

Whether the stage-two factors would sharpen if the shared direction were projected out before factoring. The factor analysis double-centres, which removes the mean but not the shared direction's within-adapter variation.

Whether the register shift is accompanied by any Big Five change: the shared-direction generations were not judged.

The shared LoRA-A drifts about ten per cent of its norm across stage two, against 1.5 per cent across stage one, which weakens the stage-two comparison relative to the stage-one one (adapter effect and drift).

Sources for this section
  • qwen35/analysis/stage2_structure.json
  • qwen35/results/fa_qwen35_stage2.json
  • qwen35/analysis/fulloct_geometry.json
  • qwen35/analysis/s2mean_steer_stats.json
  • qwen35/phase10_runs/steer_results_s2mean.json
  • qwen35/phase10_runs/steer_results_s2balanced.json
  • qwen35/results/gram_stage2.npz
  • qwen35/analysis/spider.json
  • Prose and caveats follow stage-two-structure and stage-two-shared-direction.