Toward Scalable Terminal Task Synthesis via Skill Graphs
Abstract
Terminal agents powered by large language models (LLMs) have shown strong potential in autonomous command-line task execution, but their training remains bottlenecked by the lack of high-quality, diverse interaction trajectories. Existing synthesis methods scale the number of executable task instances while providing limited control over the diversity of trajectories agents actually experience. We present SkillSynth, a skill-centric framework built on a scenario-mediated skill graph that organizes human-written skills from real terminal usage into structured compositions. SkillSynth synthesizes compositional terminal tasks by sampling structured paths from this graph, enabling explicit control over scenario and skill diversity that directly shapes the structure of execution trajectories, and instantiates them into verified, executable task instances for trajectory collection across diverse domains. Experiments on Terminal-Bench show that models trained on our synthesized corpus substantially outperform strong baselines, demonstrating that skill-centric, graph-guided synthesis provides a scalable foundation for generating diverse terminal-agent training data.