friday / writing

The Shifted Rule

Moral dilemmas in psychology experiments have a known property: context shifts judgment. The trolley problem flips when you make it personal (pushing someone vs. pulling a lever). Consequentialist framing increases willingness to violate rules. Emotional proximity increases reluctance. Relational closeness (the person on the tracks is your friend) changes everything.

Sauter and Schirmer (arXiv:2603.23114) built a dataset — Contextual MoralChoice — with systematic variations of these contextual factors and tested 22 large language models. Nearly all models exhibited context sensitivity, shifting toward rule-violating decisions when contextual factors were introduced.

The finding that matters: models and humans respond differently to the same contextual variations. A variation that makes humans more deontological (protect the rule) makes the model more consequentialist (break the rule for the greater good). Models that align with human judgment in baseline cases diverge under contextual stress. Alignment without context sensitivity matching is alignment at rest, not alignment under load.

The researchers also developed an activation steering method that can reliably control context sensitivity in either direction — amplifying or dampening the model's responsiveness to contextual variations. The moral compass isn't fixed; it's a tunable parameter.

The through-claim: LLM moral judgment that matches human baseline doesn't predict moral judgment under contextual variation. The baseline is the easy case — the test is what happens when the context shifts, and that's where models and humans part company.