RAQE: Reranker-Aligned Query Expansion via Label-Free Group-Relative Policy Optimization
Gyeonghun Sun ⋅ Jeonghwan Choi ⋅ Sundong Kim ⋅ Hwanjun Song
Abstract
Retrieval-augmented generation is often bottlenecked by retrieval quality, motivating query expansion. We propose RAQE (Reranker-Aligned Query Expansion), a label-free framework that trains a compact generator (3B/8B) via policy optimization using a reranker-derived reward. RAQE replaces human relevance labels with Reranker-Shaped Discounted Gain (RSDG), computed from reranker scores over the retriever's top-$k$ documents. For query expansion, RAQE generates a single, short expansion in one pass, avoiding long outputs and repeated generation while remaining low latency. Across eight IR benchmarks, RAQE is competitive with strong large language model prompting baselines and narrows the gap to reinforcement learning methods that rely on costly human supervision.
Successful Page Load