Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness
Abstract
Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work has studied tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment--faithfulness conflict. We show that aligned models can systematically deviate from their inputs when handling unsafe or sensitive content without explicitly disclosing these modifications, resulting in a critical failure mode we term 'alignment-induced unfaithfulness' (AIU). Unlike capability-driven unfaithfulness, which arises from errors in knowledge or reasoning, this behavior is induced by post-training alignment mechanisms that override adherence to the input. To study this, we introduce FaithConflict, a controlled dataset designed to isolate capability--faithfulness and alignment--faithfulness conflicts, along with two complementary taxonomies: a behavioral taxonomy (B1--B8) and a chain-of-thought reasoning taxonomy (C0--C5). Across a diverse set of models, we find that AIU increases with scale, often more sharply than capability-driven unfaithfulness, exhibiting a reverse scaling law. Analysis of intermediate checkpoints shows this effect is primarily amplified during post-training, with DPO emerging as a key driver. Prompting-based mitigations fail to resolve the conflict, revealing a fundamental capability--alignment--faithfulness trilemma in the design and evaluation of LLMs.