Stage two: one thing everybody learned
Open Character Training trains each persona twice. Stage one is DPO on contrasting preference pairs. Stage two is supervised fine-tuning on transcripts the model generated of itself in character, merged into stage one at a weight of 0.25. The two stages produce completely different geometries, and the difference is the most interesting thing in this dataset.
Fifteen per cent of every adapter is the same direction
Every stage-two adapter sits at almost exactly the same angle to one direction. That is not what a set of 134 different personalities should look like. The stage-two grand mean is also orthogonal to the stage-one grand mean (cosine +0.000): the two stages carry independent random LoRA-A draws, so their shared components live in different coordinates even though each is a shared component within its own stage.
What that direction does
Take the base model, add alpha times the unit grand mean of the 134 stage-two adapters, and generate. Beside it, the matched control: the same 134 adapters with exactly 67 plus and 67 minus signs, which keeps the norm and throws away the shared component (cosine +0.084 with the grand mean). Move the slider and read both columns.
The shared direction
Sign-balanced control
At alpha 0 the model is the assistant: a markdown guide addressed to "you". Positive alpha removes the guide and the second person and puts the model inside the situation. Negative alpha goes the other way, into a longer, list-heavy advisor register with falling lexical diversity. The control does not move.
Text statistics per dose
| statistic | -4.0 | -2.0 | -1.0 | 0.0 | 1.0 | 2.0 | 4.0 |
|---|---|---|---|---|---|---|---|
| stage-two grand mean | |||||||
| first person per 1k words | 21.80 | 29.30 | 32.60 | 36.10 | 46.40 | 46.70 | 47.20 |
| second person per 1k | 34.00 | 31.80 | 27.30 | 22.80 | 17.70 | 16.90 | 12.30 |
| fraction using markdown | 1.00 | 1.00 | 0.96 | 0.96 | 0.92 | 0.83 | 0.46 |
| fraction answering in character | 0.00 | 0.04 | 0.00 | 0.00 | 0.00 | 0.08 | 0.46 |
| mean words | 322.00 | 340.00 | 330.00 | 331.00 | 311.00 | 288.00 | 274.00 |
| unique-word ratio | 0.59 | 0.61 | 0.64 | 0.65 | 0.69 | 0.71 | 0.74 |
| stage-one grand mean | |||||||
| first person per 1k words | 12.80 | 19.40 | 30.70 | 34.00 | 42.50 | 43.70 | 55.90 |
| second person per 1k | 48.00 | 44.40 | 29.10 | 25.50 | 20.40 | 19.90 | 22.30 |
| fraction using markdown | 0.38 | 0.67 | 0.75 | 1.00 | 0.88 | 0.62 | 0.04 |
| fraction answering in character | 0.00 | 0.00 | 0.00 | 0.00 | 0.04 | 0.17 | 0.67 |
| mean words | 316.00 | 344.00 | 246.00 | 306.00 | 307.00 | 302.00 | 387.00 |
| unique-word ratio | 0.39 | 0.47 | 0.68 | 0.69 | 0.67 | 0.63 | 0.19 |
| sign-balanced control | |||||||
| first person per 1k words | 40.00 | 34.90 | 41.30 | 42.80 | 39.40 | 39.90 | 41.70 |
| second person per 1k | 19.30 | 19.80 | 18.90 | 17.50 | 19.70 | 16.60 | 18.80 |
| fraction using markdown | 0.96 | 0.83 | 0.92 | 0.88 | 0.92 | 0.96 | 0.96 |
| fraction answering in character | 0.04 | 0.04 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| mean words | 313.00 | 306.00 | 314.00 | 304.00 | 308.00 | 318.00 | 324.00 |
| unique-word ratio | 0.67 | 0.68 | 0.67 | 0.67 | 0.67 | 0.67 | 0.67 |
These generations were not judged on the Big Five. What is measured here is register: pronoun rates, markdown use, and whether the answer opens in the first person as the character. Each run's alpha = 0 row is the unmodified base model, and they differ from each other by GPU non-determinism alone; differences smaller than that spread should not be read. A deployed persona carries roughly alpha 0.4 of this direction, so the alpha 4 excerpts are about ten times a persona's dose. The full account.
What is left is nearly isotropic — and still stage one's arrangement
| participation ratio of the centred spectrum | 27.6 | 111.9 |
| components needed for half the variance | 20 | 52 |
| standard deviation of centred off-diagonal cosines | 0.1672 | 0.0373 |
Left column stage one, right column stage two. Out of a possible 112 of 133, the stage-two residual is spread over almost every direction available to it.
And yet
The centred off-diagonal cosines of the two stages correlate at 0.812. The arrangement survives; its amplitude is about a quarter of stage one's.
Trait by trait the two stages are orthogonal: same-trait cosine 0.00017, cross-trait 0.00002. Yet a stage-one adapter's nearest stage-two adapter is its own trait for 42 of 134 (mean rank 9.5): the tiny residual overlap is trait-specific.
The same five factors, attenuated and rotated
The same tools were run on the stage-two Gram unchanged. Parallel analysis retained 7 factors against stage one's 9. Oblimin sums of squared loadings fall from 10.76, 8.42, 7.02, 6.81, 5.80 to 2.17, 2.00, 1.95, 1.70, 1.69.
Best one-to-one matching between the stages
| Stage-one factor | Matches | Tucker φ |
|---|---|---|
| Warmth | stage-two factor 0 | +0.778 |
| Competence | stage-two factor 3 | +0.706 |
| Fearful withdrawal | stage-two factor 1 | +0.849 |
| Arousal | stage-two factor 2 | -0.702 |
| Imagination | stage-two factor 4 | +0.730 |
A negative congruence means the sign flipped; Tucker's coefficient is sign-sensitive and a flipped factor is the same factor.
Each stage-two factor's closest Goldberg target
| Factor | Target | Tucker φ |
|---|---|---|
| factor 0 | Agreeableness (Emotional stability +0.10 second) | +0.604 |
| factor 1 | Extraversion (Emotional stability +0.23 second) | +0.466 |
| factor 2 | Emotional stability (Conscientiousness +0.36 second) | +0.417 |
| factor 3 | Conscientiousness (Agreeableness -0.31 second) | +0.457 |
| factor 4 | Intellect (Extraversion +0.11 second) | +0.518 |
All five targets are taken once, at congruences below stage one's throughout.
Bipolarity is lost
In stage one a trait and its antonym sit at a negative cosine: DPO on contrasting pairs makes opposites opposite. In stage two both poles share the introspection register, the shared direction dominates their cosine, and opposite poles end up positively correlated with each other. That is the single clearest behavioural signature of what the second stage does to the weights, and it is why the stage-one factor analysis is the one to trust. Polarity and bipolarity.
The persona is stage one wearing stage two
The deployed adapter is stage one at weight 1.0 plus stage two at 0.25. Because the two stages are orthogonal, that merge is an exact Pythagorean sum, and the norms work out so that each stage contributes about half of the persona's squared norm.
Equal norms, and the geometry is stage one's. Weight energy and behavioural content come apart here as they do everywhere else in this project: half the persona's weight change is a register direction that carries almost no trait information, and the arrangement that identifies the trait survives in the other half. Full OCT replication.
The dials tell the other half of the story: on the OCEAN dials the deployed personas move the judge's scores further than the stage-one adapters alone do, on 9 of the 10 dials. Adding the register direction makes the model act more like the character even though it adds no trait geometry. See the dials.
Further stage-two findings
What stage two installs
Open Character Training's second stage has each stage-one model generate 12,000 transcripts about itself, fine-tunes a fresh rank-64 LoRA on them, and releases the two merged at weight 0.25. The striking fact about the 134 resulting adapters is how much they share: 0.151 of every adapter's squared norm lies along one direction, the grand mean, and every adapter sits at cosine 0.389 to it with a spread of 0.031 (analysis/stage2_structure.json). Steered alone that direction turns markdown advice addressed to "you" into first-person in-character answers.
Checking that this is not an artefact of initialisation corrected the record twice. First, the 134 stage-two adapters do not each have their own random LoRA-A, as the wiki said: all were trained at sft_seed 123456 and draw the same one, still at pairwise cosine 0.9977 to 0.9981 after training against 0.0021 across seeds (analysis/lora_a_identity.json). A rank-64 delta lives in the row space its LoRA-A defines, which is why stage-two cross-trait cosines are +0.1452 within a seed and +0.0141 across seeds. The shared component survives the check: on the 15 traits put through the pipeline a second time it is 0.2017 of the squared norm at seed 0 and 0.2000 at seed 1, at cosine 0.4497 +- 0.0319 and 0.4475 +- 0.0330 (analysis/stage2_frame.json). It is a property of the recipe; only its coordinates come from the initialisation.
Second, a units error. Every steering run sets ref to 0.8078003190997738 and calls it the mean adapter norm, but it was computed as the mean of ||B @ A|| without the LoRA scaling 2.0; the real mean is 1.6157416444226869, so alpha 1 is half an adapter's worth of weight change, not one (analysis/steer_alpha_units.json). Nothing measured changes, but every published dose reads as half: a persona's dose of the shared direction is alpha 0.7788, not the 0.4 previously stated.
Where the behavioural amplification comes from
Personas move their own behavioural dial further than their stage-one adapters do, even though the persona's geometry is stage one's. Stage two adds two things
- the register direction and a trait-specific residual - so which is responsible?
For the 15 traits with a second seed, four conditions on the 24-prompt Big Five battery, all built by adding an fp32 delta into the bf16 base weights, judged blind (analysis/stage2_register_vs_residual.json; 1,464 generations):
| condition | own-factor amplification | in-character fraction | markdown fraction | |---|---|---|---| | base | - | 0.00 | 0.92 | | stage one alone | +22.33 | 0.48 | 0.54 | | stage one + the shared direction at 0.5x a persona's dose | +22.40 | 0.50 | 0.52 | | stage one + the shared direction at 1.28x a persona's dose | +24.52 | 0.54 | 0.50 | | the exact persona | +30.25 | 0.60 | 0.37 |
No. Half a persona's dose of the register buys 0.9 percent of the stage-one-to-persona gap and 1.28 times the dose buys 27.6 percent, neither significant over 15 paired traits, while the persona itself gains +7.92 (t 2.83; 11 of 15 traits move above their own stage-one adapter). The norms agree: of the persona's stage-two half, 0.629 of the Frobenius norm is the register and 1.489 the residual. Steering the register hard dominates the text, but at the strength a persona carries it, it is not what makes a persona more itself. The run also validates the mechanism: its stage-one condition reproduces the existing stage-one evaluation at +22.33 against +22.26, per-trait Pearson 0.937.
What stage two looks like in activation space
Running the same 64 prompts with each adapter as forward hooks and recording the mean residual-stream shift gives a different view of the same objects (analysis/actspace_stage2_geometry.json, layer 16, response window). In weight space the two stages are as orthogonal as possible: the same trait's deltas have cosine +0.0002 and the two grand means +0.000. In activation space the same trait's two shifts have cosine +0.530 and the two mean shifts +0.552. Orthogonality in weight coordinates is a fact about the parameterisation, not about what the stages do to the model - which is why two directions orthogonal in coordinates produce the same register shift when steered.
The shared component is larger in activations and ordered the same way: cosine to the mean shift 0.842 +- 0.070 for stage two against 0.701 +- 0.079 for stage one. The persona's activation geometry is stage one's (same-trait cosine +0.988, all 134 nearest neighbours correct, arrangement correlation +0.993). And a stage-two LoRA alone still knows its trait: its own constitution's prompted activation vector is its nearest of 134 for 50 of 134, against 43 for stage one and 55 for the persona.
How to read its factors
Parallel analysis retains 7 factors for stage two at N = 1528, against 9 for stage one. Recomputing the factor analysis on a Gram with the shared direction projected out of the deltas exactly changes nothing: the loading matrices match at Tucker 0.990 to 0.99996, the retained count is still 7, and the top reduced eigenvalues move from 3.97, 2.98, 2.46, 2.24, 2.05 to 3.97, 2.98, 2.46, 2.20, 1.99 (analysis/stage2_factors_choice.json). The factor analysis already double-centres the correlation matrix, which does the same job. Nor is k = 7 more interpretable: its seven factors take only six distinct Big Five targets, one of them the general evaluative axis, Conscientiousness is claimed twice, and its three extra factors decompose the k = 5 Emotional Stability factor. Present k = 5 on the ordinary stage-two Gram, noting the retained 7 and that at k = 7 a clean orderliness factor (systematic, organized, prompt, neat against quiet, untalkative, withdrawn) separates out of the muddled k = 5 Conscientiousness factor.
Whether the register is persona-specific
The control that had never been run. The whole introspection stage, rerun with a trait-free constitution - a helpful, honest assistant with no distinctive personality - on the plain base model with no stage-one adapter anywhere, five times with independent generation seeds, at the zoo's own sft_seed so the LoRA-A frames match (checked afterwards: 0.9982).
| | five neutral adapters | the 134 zoo stage-two adapters | |---|---|---| | \|dW\| | 6.4976 +- 0.0082 | 6.4643 +- 0.2899 | | cosine with the stage-two grand mean | 0.3035 +- 0.0007 | 0.3893 +- 0.0311 | | cosine with an arbitrary trait adapter | +0.1179 +- 0.0225 | +0.1452 +- 0.0359 |
Mostly generic. An adapter trained on a constitution saying the character has no character still reaches 78 percent of the shared direction the 134 personas sit on, and 81 percent of their mutual similarity. What stage two installs is mostly "narrate yourself in the first person"; the trait is close to incidental to that lesson. Not entirely: 0.3035 is below the zoo's mean minus two standard deviations, and the five neutral runs agree to within 0.002, so conditioning on a trait adds about a fifth of the shared component.
Two things fell out. The five sit at +0.5633 with each other against the zoo's +0.1452, and still +0.5189 after the shared direction is removed: given the same constitution and base, this recipe is close to deterministic. And every one has the same nearest neighbour among the 134 - unemotional, at +0.167 - with its residual profile correlating with unemotional's own at r 0.86. A constitution saying "neither warm nor cold" and "under pressure nothing changes" landed on the trait word for exactly that, without being told.
What remains open
The register experiment rests on 15 traits and one judge. The activation-space comparison runs a stage-two LoRA alone on a base it was not trained on. The neutral control trains on the plain base while the 134 trained on 134 different DPO-merged bases; the argument that this matters little is that the shared direction is already at 0.389 across all 134 differently perturbed bases. And nothing here tests whether any of it survives a change of base model or recipe.
From companion/content/stage2_findings.md.
Stage-two exploration
| date | 2026-09-08 |
| what | Four experiments on OCT stage two: register versus residual, a trait-free control, activation space, and the number of factors. Detail in the four files named below. |
| frame.file | analysis/stage2_frame.json |
| frame.lora_a_shared_within_a_seed_cos | 0.9979 |
| frame.lora_a_across_seeds_cos | 0.002142 |
| frame.shared_share_seed0_same15 | 0.2017 |
| frame.shared_share_seed1_15 | 0.2 |
| frame.cos_to_mean_seed0_same15 | 0.4497 |
| frame.cos_to_mean_seed1_15 | 0.4475 |
| frame.reading | the shared direction reproduces in an independent LoRA-A frame |
| register_vs_residual.file | analysis/stage2_register_vs_residual.json |
| register_vs_residual.amplification_cond_b | 22.33 |
| register_vs_residual.amplification_cond_c | 22.4 |
| register_vs_residual.amplification_cond_d | 24.52 |
| register_vs_residual.amplification_cond_e | 30.25 |
| register_vs_residual.share_of_gap_cond_c | 0.009326 |
| register_vs_residual.share_of_gap_cond_d | 0.2761 |
| register_vs_residual.first_person_per_k_cond_a | 41.2 |
| register_vs_residual.markdown_frac_cond_a | 0.92 |
| register_vs_residual.first_person_per_k_cond_b | 44.8 |
| register_vs_residual.markdown_frac_cond_b | 0.54 |
| register_vs_residual.first_person_per_k_cond_c | 45.3 |
| register_vs_residual.markdown_frac_cond_c | 0.52 |
| register_vs_residual.first_person_per_k_cond_d | 45 |
| register_vs_residual.markdown_frac_cond_d | 0.5 |
| register_vs_residual.first_person_per_k_cond_e | 51.1 |
| register_vs_residual.markdown_frac_cond_e | 0.37 |
| neutral_control.file | analysis/stage2_neutral_control.json |
| neutral_control.n_neutral | 5 |
| neutral_control.neutral_cos_with_grand_mean_mean | 0.3035 |
| neutral_control.neutral_cos_with_grand_mean_sd | 0.0006898 |
| neutral_control.zoo_cos_with_grand_mean_mean | 0.3893 |
| neutral_control.neutral_x_neutral_mean | 0.5633 |
| neutral_control.neutral_x_zoo_mean | 0.1179 |
| neutral_control.zoo_offdiag_mean | 0.1452 |
| neutral_control.neutral_norm_mean | 6.498 |
| neutral_control.zoo_norm_mean | 6.464 |
| activation_space.file | analysis/actspace_stage2_geometry.json |
| activation_space.primary_layer | 16 |
| activation_space.act_cos_to_mean_stage1 | 0.7006 |
| activation_space.act_shared_share_stage1 | 0.4786 |
| activation_space.act_cos_own_prompt_stage1 | 0.6009 |
| activation_space.act_rank1_stage1 | 43 |
| activation_space.act_cos_to_mean_stage2 | 0.8419 |
| activation_space.act_shared_share_stage2 | 0.6693 |
| activation_space.act_cos_own_prompt_stage2 | 0.7623 |
| activation_space.act_rank1_stage2 | 50 |
| activation_space.act_cos_to_mean_persona | 0.7052 |
| activation_space.act_shared_share_persona | 0.4856 |
| activation_space.act_cos_own_prompt_persona | 0.6413 |
| activation_space.act_rank1_persona | 55 |
| activation_space.act_vs_weight_stage1_stage1 | 0.8652 |
| activation_space.act_vs_weight_stage2_stage1 | 0.7869 |
| activation_space.act_vs_weight_stage1_stage2 | 0.7132 |
| activation_space.act_vs_weight_stage2_stage2 | 0.7325 |
| activation_space.act_vs_weight_stage1_persona | 0.8669 |
| activation_space.act_vs_weight_stage2_persona | 0.8002 |
| activation_space.act_cos_mean_shifts_persona_x_stage1 | 0.9929 |
| activation_space.act_same_trait_persona_x_stage1 | 0.988 |
| activation_space.act_cos_mean_shifts_persona_x_stage2 | 0.6265 |
| activation_space.act_same_trait_persona_x_stage2 | 0.5844 |
| activation_space.act_cos_mean_shifts_stage1_x_stage2 | 0.5516 |
| activation_space.act_same_trait_stage1_x_stage2 | 0.5299 |
| factors.file | analysis/stage2_factors_choice.json |
| factors.stage2_chosen_k | 7 |
| factors.stage2_noshared_chosen_k | 7 |
| factors.stage1_chosen_k | 9 |
| factors.gram_trace_share_removed | 0.1528 |
| factors.gram_offdiag_cos_before | 0.1452 |
| factors.gram_offdiag_cos_after | -0.007492 |
| factors.gram_corr_before_after | 0.8791 |
| factors.targets_taken_stage2|centred_k5 | 5 |
| factors.mean_abs_congruence_stage2|centred_k5 | 0.4923 |
| factors.targets_taken_stage2|centred_k7 | 6 |
| factors.mean_abs_congruence_stage2|centred_k7 | 0.4538 |
| factors.targets_taken_stage2_noshared|centred_k5 | 5 |
| factors.mean_abs_congruence_stage2_noshared|centred_k5 | 0.5031 |
| factors.targets_taken_stage2_noshared|centred_k7 | 6 |
| factors.mean_abs_congruence_stage2_noshared|centred_k7 | 0.4588 |
qwen35/analysis/stage2_exploration.json
What stage two does not establish
Whether the shared direction is specific to the introspection recipe or would appear after any supervised fine-tuning on self-generated text: no control SFT arm exists.
Whether the stage-two factors would sharpen if the shared direction were projected out before factoring. The factor analysis double-centres, which removes the mean but not the shared direction's within-adapter variation.
Whether the register shift is accompanied by any Big Five change: the shared-direction generations were not judged.
The shared LoRA-A drifts about ten per cent of its norm across stage two, against 1.5 per cent across stage one, which weakens the stage-two comparison relative to the stage-one one (adapter effect and drift).
Sources for this section
qwen35/analysis/stage2_structure.jsonqwen35/results/fa_qwen35_stage2.jsonqwen35/analysis/fulloct_geometry.jsonqwen35/analysis/s2mean_steer_stats.jsonqwen35/phase10_runs/steer_results_s2mean.jsonqwen35/phase10_runs/steer_results_s2balanced.jsonqwen35/results/gram_stage2.npzqwen35/analysis/spider.json- Prose and caveats follow stage-two-structure and stage-two-shared-direction.