Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification
Abstract
Real-world visual question answering (VQA) may be \emph{context-dependent}: an image-question pair could be under-specified, such that the correct answer depends on external information that is not observable in the image. In such cases, directly answering can lead to confident but incorrect predictions. We propose CoA (Clarify-or-Answer), an ask-or-answer agent that separately models the decision to ask or answer, and what to ask if needed. CoA first determines whether clarification is necessary; if so, it generates one or more clarification questions to iteratively enrich the context until ambiguity is resolved. To validate this framework, we introduce \dataset, a dataset of ambiguous VQA questions along with a contrast set of non-ambiguous questions. We further introduce \method (Clarification Reasoning), a reinforcement learning approach that optimizes clarification question generation with multiple reward signals encouraging well-formed, focused, non-trivial questions that resolve ambiguity. Across three VLMs and three datasets, CoA achieves consistent improvements at both the module and system levels, with a single clarification turn providing significant gains and recursive multi-turn execution further boosting accuracy in some settings beyond baselines.