Rule vs. Consequence: Dissociable Internal Representations of Moral Reasoning in LLMs
Eugenie Shi
Abstract
Large language models are increasingly deployed in morally and legally sensitive contexts, yet little is known about how they internally represent distinctions between moral reasoning styles. We construct a 2$\times$2 factorial dataset crossing reasoning style (rule-based vs. consequence-based framing) with judgment polarity (permissive vs. prohibitive), drawing from three established moral reasoning benchmarks (ETHICS, Moral Stories, Social Chemistry 101) and evaluating on a cleanroom test set derived from LegalBench and hand-written scenarios, including conflict pairs where the two framings predict opposite labels. Using linear probing and activation steering across three instruction-tuned LLMs from different developers, we identify two orthogonal directions in the residual stream: Dir-A encodes moral reasoning style ($r > 0.87$ with style labels, $r < 0.03$ with polarity labels) and Dir-B encodes judgment polarity ($r > 0.85$ with polarity labels, $r < 0.06$ with style labels), with cross-correlation $|r| < 0.04$ between the two directions. This double dissociation means each direction captures one factor while remaining blind to the other, and it replicates consistently across Qwen2.5-7B-Instruct, Mistral-7B-Instruct-v0.3, and Llama-3.1-8B-Instruct at 34 to 47 percent network depth. The framing-sensitive direction is robust across layers: all ten layers tested in the range 6 to 15 show significant framing selectivity ($p < 0.0001$), ruling out a narrow layer-specific artifact. Activation steering with Dir-A produces large causal shifts (Cohen's $d = 1.03$ to $1.66$, all $p < 0.0001$), and on conflict scenarios steering selectively corrects rule-framed errors by 30 percentage points while leaving consequence-framed accuracy unchanged. Cross-domain generalization achieves 0.967 mean accuracy across three of four benchmark sources, and the activation probe exceeds an $n$-gram lexical baseline on novel-template evaluation (0.594 vs. 0.438), providing evidence of partial transfer beyond surface features.
Successful Page Load