Personality in weight space 134 trait adapters · Qwen3.5-4B

Methods

Five short pages, one per pillar of the work. Each one summarises what was done and hands off to the wiki, which carries every number with the file and JSON key it came from.

ConstructionHow 140 trait words became 134 adaptersGeometryThe exact Gram, the factor analysis, and the controlsBehaviourSteering, judging and the ceiling on bothActivation spaceWhether prompting and training move the model the same wayScoring training dataReading a dataset against a direction in one backward pass
Vocabulary

Three terms worth fixing

Stage one and stage two. Stage one is the DPO adapter trained on contrasting preference pairs. Stage two is the introspection SFT adapter trained on self-generated in-character transcripts. The deployed persona is stage one at 1.0 plus stage two at 0.25. The two stages are trained from independent random LoRA-A draws, so they are orthogonal in coordinates even where their arrangements agree.

Intellect, not Openness. The fifth factor is called Intellect here, following Goldberg's own label for the 100-marker set the traits were drawn from. Why the fifth factor has two names.

Loadings and coordinates are different objects. A loading is a trait's weight on an oblique factor in the pattern matrix. A chart coordinate is an inner product with an orthonormal basis of the five factors' span, so it says where the adapter is, not what it loads on. The scatter uses coordinates; the loading panels use loadings. The glossary.

Sources for this section
  • qwen35/results/fa_qwen35.json
  • qwen35/analysis/actspace_geometry.json
  • qwen35/analysis/actspace_adapters_geometry.json
  • qwen35/analysis/nxn_summary.json
  • qwen35/analysis/scree_null_matched.json
  • qwen35/analysis/crossseed_arms.json
  • Every claim on these pages is a summary of a wiki page, linked in place.