Capability Provenance in Language Models: A Case Study in Social Reasoning
Glenn Matlin ⋅ Chandreyi Chakraborty ⋅ Saehee Eom ⋅ Mika Okamoto ⋅ Rayan Castilla ⋅ Louis Jaburi ⋅ Alvin Deng ⋅ Taywon Min ⋅ Lucia Quirke ⋅ Stella Biderman ⋅ Mark Riedl
Abstract
Where do social-reasoning capabilities in language models come from in pretraining data? While prior work has studied document-level attribution, reasoning behaviors --- especially social reasoning --- are not well-understood, and may arise from broader, more distributed corpus structure. We study this question in the OLMo3 ecosystem using gradient-based training-data attribution (TrackStar) over a stratified working set drawn from the de-duplicated Dolma3 training corpus. Rather than interpreting individual document scores, we aggregate influence over the WebOrganizer topic $\times$ format taxonomy and compare attribution distributions across a $2 \times 2$ contrastive benchmark design: SocialIQA and MMLU Social Sciences versus GSM8K and MMLU-STEM. Social reasoning and mathematical reasoning draw on qualitatively different corpus regions---social-reasoning influence concentrates in bins that mathematical reasoning does not depend on, and vice versa---in patterns that are consistent across the taxonomy. Targeted machine unlearning further validates these associations: selectively forgetting high-attribution topic bins degrades the corresponding benchmark while leaving contrastive controls largely intact. Together, our results demonstrate a corpus-level framework for capability provenance, showing that social reasoning in language models arises from distinct, interpretable regions of pretraining data. We open-source all code, data artifacts, influence scores, and checkpoints.
Successful Page Load