FitText: Evolving Agent Tool Ecologies via Memetic Retrieval
Abstract
A fundamental semantic gap separates how users describe tasks from how tools are documented, and as API ecosystems grow to tens of thousands of endpoints, static upfront retrieval from the initial query cannot bridge it: the agent's understanding of what it needs evolves, but its retrieval does not. We argue that tool retrieval must be dynamic, embedded directly in the reasoning loop, and allowed to sharpen as the agent's comprehension of the task deepens. We introduce FitText, a training-free framework that generates natural-language pseudo-tool descriptions as mutable retrieval probes, iteratively refines them against the tool database using retrieval feedback, explores diverse alternative hypotheses through stochastic generation, and evolves the strongest candidates via selection pressure and a tool memory that steers away from previously hypothesized tool regions. On ToolRet (43k tools, 4 domains), FitText improves average retrieval rank from 6.88 to 2.88; on StableToolBench (16,464 APIs), it achieves a 0.73 average pass rate, a 24-point absolute gain over static query retrieval, with consistent improvements across multiple base models.