Learning to Comply: Workflow-Grounded Environment Generation for Training Procedurally Compliant Agents
Abstract
Recent advances in AI agents have expanded their applicability beyond simple task completion to increasingly complex real-world settings. In such settings, the challenge extends beyond solving difficult tasks; it also requires following the behavioral policies or operational procedures that govern how those tasks should be carried out. However, how to systematically train agents for such procedural compliance remains underexplored, especially along two dimensions: (1) how to construct environments for compliance training, and (2) how to provide learning signals that reward compliant behavior. In this work, we introduce a scalable, automated, and verifiable pipeline for jointly generating executable tool-use environments and tasks, each paired with a ground-truth execution path. We further investigate diverse reward design strategies inspired by sequence-matching algorithms to incentivize procedurally compliant behavior, with rewards defined at both the step and trajectory levels. In experiments across diverse benchmarks requiring procedural compliance in tool-use tasks, we find that step-level rewards improve credit assignment by enabling finer-grained advantage estimation in PPO, while trajectory-level rewards, as used in GRPO, yield strong group-relative learning signals when coupled with our environment-specific reward design. Collectively, our results establish that scalable environment construction and compliance-aware reward design are both essential for training agents that are procedurally grounded.