Tool-Creating LLM Agents Gain Little from Keeping Their Tools
Marek Šuppa ⋅ Jaroslav Kopcan
Abstract
Tool-creating LLM agents are widely reported to benefit from accumulating reusable tool libraries. We test this claim with a controlled ablation: create a tool, then throw it away. Across two benchmarks (BigCodeBench-Hard, 148 tasks; $\tau^2$-airline, 50 tasks), four models, and three retrieval variants including embedding-based retrieval, creating and \emph{discarding} tools consistently matches or exceeds creating and \emph{keeping} them. On BigCodeBench-Hard, all eight ablation conditions cluster within a 2pp band. On $\tau^2$-airline, a wrapper-matched decomposition reveals that the raw +18pp Direct-to-Discard gain splits into +10pp from the agent harness itself, +6pp from a simple ``use your tools'' instruction, and only +2pp from the tool-creation structure. These results suggest that tool creation helps agents engage with available capabilities, not build reusable memory---and that agent-framework details can inflate reported gains by 10pp or more.
Successful Page Load