Learning Steerable Clarification Policies with Collaborative Self-play
Abstract
To handle ambiguous queries, AI assistants must decide (a) when to guess the user's intent and answer directly, (b) when to enumerate and respond to multiple plausible intents, and (c) when to ask a clarifying question. Importantly, the right decision depends on contextual factors such as user preferences or modality. For instance, enumerating several possible interpretations can be cumbersome on small screens or voice-based settings. In this work, we propose training steerable policies for managing such uncertainty via self-play. Given two agents, one simulating a user and the other an AI assistant, we generate conversations where the user issues a potentially ambiguous query, and the assistant must choose how to respond. The model receives as input the numerical cost of asking a clarification question, and of generating each word, and is trained to select the action that maximizes its final reward: accuracy penalized by these costs. We apply Reinforced Self-Training (ReST) to optimize this objective and show that it yields a steerable policy that adapts its behavior based on the provided costs, improving both reward and accuracy. Moreover, the resulting policy generalizes to numerical cost values unobserved during training.