Bridging Databases and Documents: Data-Algorithm Co-Design for Hybrid Question Answering
Abstract
Large Language Models (LLMs) increasingly rely on multi-tool orchestration to resolve complex enterprise queries. However, answering hybrid questions requires tightly coupled reasoning steps between structured data retrieval and unstructured semantic search, where a single execution error triggers an irreversible cascade. Training models to navigate these strict cross-modal pipelines via standard Reinforcement Learning (RL) is notoriously difficult. Because successful multi-hop trajectories are rare, unguided exploration leads to severe data waste and sparse rewards. While recent algorithmic advancements improve RL optimization dynamics, they still fundamentally rely on the highly inefficient process of discovering successful trajectories from scratch. In this paper, we advocate for a data-algorithm co-design paradigm to overcome this exploration bottleneck. First, we introduce HyDRA-Bench, a novel environment simulating realistic hybrid reasoning where agents must seamlessly interleave SQL and search. Second, we propose TRACE-GRPO, an adaptive RL framework that tightly integrates the data generation pipeline with policy optimization. By converting intermediate reasoning traces generated during dataset synthesis into executable, action-level hints that serve as proxy goals, TRACE-GRPO transforms sparse unguided rollouts into a dense, highly efficient curriculum. This approach not only stabilizes optimization by preventing gradient vanishing but also drives ``in-context to in-weights'' skill elicitation. Empirical results demonstrate that TRACE-GRPO drastically accelerates convergence, enabling small open-weights models to rival top-tier proprietary models on hybrid reasoning, while exhibiting robust generalization to broader multi-hop text environments like MuSiQue.