From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent advances in reasoning-oriented systems by enabling large-scale optimization over automatically checkable outcomes. However, this paradigm is most effective in domains such as mathematics and coding, where correctness can be deterministically verified, and becomes much less reliable in open-ended tasks that require nuanced judgment, incomplete evidence, and subjective trade-offs. Existing attempts to extend reinforcement learning to such settings typically replace ground-truth rewards with proxy signals inferred directly or indirectly by the actor itself, resulting in biased supervision and a tight coupling between the capabilities of the actor and the judge. In this work, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a new training paradigm that bridges this gap by transforming open-domain tasks into verifiable proxy tasks through self-play. Specifically, we formulate the proxy task as a multi-agent adversarial game governed by simple rule-based evaluation, so that rewards emerge automatically from agent interactions without requiring human annotations, external verifiers, or additional datasets. We instantiate this framework with a game analogous to Who Is the Undercover, in which agents are assigned asymmetric information roles and must produce statements and vote strategically under partial observability. Although the proxy task is fully verifiable, it preserves substantial capability overlap with the target tasks, enabling competitive training to transfer to open-domain settings. We apply RLSVR to text summarization, creative writing, and mathematical reasoning, and experimental results show that it consistently outperforms existing self-evolution methods. These findings suggest that self-play can serve as an effective mechanism for scalable reinforcement learning in domains where direct reward verification is otherwise difficult.