Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
Abstract
How can we build models whose post-trained capabilities survive subsequent fine-tuning? Rather than focusing on downstream forgetting interventions alone, we study how upstream training choices could make models inherently more robust to later adaptation. We investigate this question in a controlled three-stage language-model pipeline: pretraining, post-training to acquire a target capability, and downstream fine-tuning on a new objective. Across 135M and 1B parameter models, we find that immediate post-training performance does not reliably predict downstream retention: training choices that seem equivalent at the end of post-training can produce substantially different forgetting after subsequent fine-tuning. In particular, mixing post-training data into pretraining consistently improves the frontier between retained upstream capability and downstream fine-tuning loss, including in regimes where it has little effect on immediate post-training loss. In compute-matched experiments, mixing and dedicated post-training serve different functions: post-training drives immediate specialization, while mixing improves robustness to later forgetting. Replay and dropout during post-training provide additional, complementary gains. These results position robustness to subsequent fine-tuning as a distinct objective of upstream model development and show that how a model learns a capability shapes how robustly it is retained.