Fine-Tuning Was Not Broken but Diffuse: Precise Knowledge Editing via Projected Error Signals
Abstract
Transformer-based large language models (LLMs), trained via backpropagation on massive textual corpora, encode substantial factual knowledge, yet may also retain false or outdated associations. Prior work on factual editing in LLMs shows that fine-tuning (FT) often induces a repelling effect, degrading generalization and overwriting non-targeted knowledge, motivating alternative approaches such as rank-1 optimizations. In this work, we revisit knowledge editing in LLMs through FT by analyzing the activations and Vector-Jacobian Products (VJPs), the error signals produced during backpropagation when learning new facts. We show that each layer and token induces its own error signals, and that attempts to edit a specific association can inadvertently encourage the model to memorize the entire prompt in a diffuse or inefficient manner. Building on this analysis, we introduce Precise-FT (PFT), a method that constructs gradient updates from a small set of projected error signals most strongly associated with the learned fact, achieving state-of-the-art performance across multi-knowledge editing benchmarks and LLMs.