Synthetic Sandbox for Training ML Engineering Agents
Abstract
As large language model (LLM) agents advance beyond software engineering (SWE) toward machine learning engineering (MLE), verifying agent behavior becomes orders of magnitude more expensive: while SWE tasks can be verified via fast-executing unit tests, MLE verification requires running full ML pipelines---data preprocessing, model training, and metric evaluation---on large datasets at each rollout step, rendering trajectory-wise on-policy reinforcement learning (RL) prohibitively slow. Existing approaches retreat to supervised fine-tuning (SFT) or offline proxy rewards, sacrificing the exploration and generalization benefits of on-policy RL. We argue that sandbox data size is one of the primary sources of this bottleneck. Based on this insight, we introduce SandMLE, a framework that generates diverse, verifiable synthetic MLE environments from seed tasks, preserving the structural and mathematical complexity of real-world problems while constraining datasets to micro-scale (50–200 samples). SandMLE reduces execution time by over 13×, enabling large-scale, on-policy trajectory-wise RL for the first time in the MLE domain. On MLE-bench-lite, SandMLE yields a 20.3\% to 66.9\% relative improvement in medal rate over the SFT baselines across the Qwen3-8B, 14B, and 30B-A3B. Furthermore, the trained policy generalizes across unseen agentic scaffolds, achieving up to 32.4\% relative improvement in HumanRank score on MLE-Dojo.