RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Abstract
Most reward models for visual generation compress rich human judgments into a single unexplained score, discarding the structured reasoning that underlies preference. We show that equipping reward models with explicit, multi-dimensional chain-of-thought critiques transforms them from passive evaluators into versatile optimization interfaces, unlocking improvements in two complementary spaces: - Parameter space, where structured rationales provide semantically grounded rewards for reinforcement learning, and - Prompt space, where a Generate–Critique–Refine loop translates critiques into targeted prompt revisions at test time—without parameter updates. To build such a model, we introduce Preference-Anchored Rationalization (PARROT), a variational framework that treats rationales as latent variables and derives an ELBO whose three terms map directly onto a scalable data-synthesis pipeline: preference-anchored rationale generation, predictive consistency filtering, and foresight distillation.The resulting model, RationalRewards (8B), achieves state-of-the-art preference prediction among open-source reward models—competitive with Gemini-2.5-Pro—while using 10–20× less training data than comparable baselines. As an RL reward, it consistently improves text-to-image and image-editing generators beyond scalar baselines. Notably, its test-time Generate–Critique–Refine loop matches or exceeds RL-based fine-tuning on several benchmarks, suggesting that structured critiques can unlock latent generator capabilities that suboptimal prompts fail to elicit.