Detecting and Suppressing Reward Hacking with Gradient Fingerprints
Abstract
Reinforcement learning with verifiable rewards (RLVR) imposes no constraints on intermediate reasoning. This leaves it susceptible to reward hacking, where models exploit loopholes in the reward function to achieve high scores without solving the intended task. These reward-hacking behaviors are often implicit, as the chain-of-thought (CoT) may still appear plausible, limiting the effectiveness of text-based monitors. We propose GRIFT (Gradient Fingerprint), a method for detecting reward hacking behavior through models' internal computation. Given a prompt and a CoT generated from a model, GRIFT encodes the CoT into a compact representation, called a gradient fingerprint, derived from gradients of the CoT conditioned on the prompt. This representation is computed efficiently using lightweight adapters on selected layers and further compressed via random projection. The proposed method enables accurate reward hacking detection with minimal supervision. Across verifiable reasoning benchmarks spanning math, code, and logical reasoning, GRIFT substantially outperforms strong baselines, including CoT Monitor and TRACE, achieving over 25% relative improvement in detection performance. Moreover, integrating GRIFT into a rejection fine-tuning pipeline for sample selection reduces reward hacking and improves intended task performance. Overall, our results demonstrate the effectiveness of gradient-level representations for detecting reward hacking and highlight their promise for assessing reasoning quality beyond surface-level text.