Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
Abstract
Mechanistic interpretability has made it possible to localize circuits underlying specific behaviors in language models, but existing methods are expensive, model-specific, and difficult to scale to larger architectures. We introduce Differentiable Faithfulness Alignment (DFA), a framework for transferring circuit information from a smaller source model to a larger target model through a learned differentiable alignment. DFA maps source-model node importance scores into the target model and optimizes this mapping for causal faithfulness via a soft intervention objective, avoiding full circuit discovery on the target model. We evaluate DFA on Llama-3 and Qwen-2.5 across six tasks spanning factual retrieval, multiple-choice reasoning, and arithmetic. The strongest results occur on LLaMA-3-1B -> 3B, where aligned circuits are often competitive with direct target-model node attribution and zero-shot transfer remains effective. Recovery is weaker for larger source--target gaps and substantially worse on Qwen-2.5, suggesting that transfer becomes harder as architectural and scaling differences increase. Overall, DFA consistently improves over simple baselines and, in some settings, recovers target-model circuits with faithfulness comparable to direct attribution. These results suggest that smaller models can provide useful mechanistic priors for larger ones, while highlighting both the promise and the limits of node-level cross-model circuit alignment.