Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
Abstract
Reinforcement learning from human feedback (RLHF) is vulnerable to reward hacking: policies can exploit spurious correlations in learned reward models to increase proxy reward while degrading true alignment. Existing defenses are largely static. They regularize optimization, improve the reward model, or target known biases, but do not explicitly detect exploitation as it emerges during training. We propose \textbf{Adversarial Reward Auditing (ARA)}, a two-stage framework for detecting and mitigating reward hacking. In Stage 1, a Hacker policy searches for vulnerabilities in a frozen reward model while an Auditor learns to detect exploitative responses from the reward model's internal representations. In Stage 2, the trained Auditor guides RLHF by downweighting rewards for responses likely to be exploitative. Across sycophancy, length bias, and code gaming, ARA yields the strongest alignment--utility tradeoff among competitive baselines, reducing exploitation while preserving or improving task performance. We also find cross-domain transfer in both attack and defense: a Hacker trained on code gaming increases sycophancy despite receiving no direct reward for that behavior, while an Auditor trained in one domain helps suppress exploits in others. These results suggest that reward hacking is a dynamic, transferable phenomenon and that auditing optimization in reward-model representation space is an effective way to control it.