Behaviour
Steering, judging and the ceiling on both
A weight-space direction is turned into behaviour by adding alpha times the unit direction times a reference norm (0.81; the final audit found it omitted the LoRA scale of 2, so alpha 1 is half of one stage-one adapter's Frobenius norm and alpha 2 is one adapter) to the base model, then generating 512 tokens greedily on 24 fixed prompts with thinking disabled. An LLM judge scores each answer blind on the five scales from 1 to 7.
Two defects shape every behavioural number here. First, Qwen3.5's chat template defaults reasoning on: three steering runs that omitted the flag were silently invalidated, regenerated at 512 tokens, and one built page was withdrawn. Second, about 59 per cent of responses in the corrected corpus end mid-sentence at the token cap, so no claim is made about how a response concludes, and looping responses are counted as damage before any effect is read.
The judge's own reliability was measured rather than assumed, and it caps what can be claimed. Against that ceiling: most factor directions move their own scale, additivity fails at about half the predicted effect, and one factor moves everything at once. The honest summary is that the geometry replicates and the behaviour does not add up.