Prescriptive Scaling Laws for Data Constrained Training
Justin Lovelace ⋅ Christian Belardi ⋅ Srivatsa Kundurthy ⋅ Shriya Sudhakar ⋅ Kilian Weinberger
Abstract
Training compute is increasingly outpacing the availability of high-quality data, shifting the central challenge from optimal compute allocation to extracting maximum value from limited data. The widely adopted Chinchilla scaling law assumes every training token is unique, and existing extensions for data repetition fail to accurately model overfitting---limiting their ability to guide pretraining decisions in data-constrained regimes. We conduct a large-scale repetition study---over 300 runs spanning 15M--1B parameters, 50M--6B unique tokens, two weight decay strengths, and up to 16 epochs. We model the excess loss under repetition with an additive overfitting penalty and find that repetition damage is superlinear: each additional epoch inflicts more damage than the last. Even a one-parameter additive form substantially outperforms prior work on both our data and the Muennighoff et al. (2023) data. The superlinear penalty yields qualitatively new compute-optimal allocation advice: beyond a data-dependent threshold, further repetition is counterproductive and compute is better spent on model capacity. We validate this prescriptively, training each law's recommended configuration and showing that ours achieves the strongest performance in both perplexity and downstream evaluation. Finally, because our one-parameter form isolates overfitting in a single coefficient, it enables direct comparison across training configurations; as a case study, we show that strong weight decay ($\lambda=1.0$) reduces this coefficient by approximately 70\%, providing a scaling-law explanation for recent findings that optimal weight decay in data-constrained regimes is an order of magnitude larger than standard practice.
Successful Page Load