RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations
Yuxuan Zhang ⋅ Cong Wei ⋅ Penghui Du ⋅ Bo Li ⋅ Huaisong Zhang ⋅ Songcheng Cai ⋅ Yubo Wang ⋅ Dongfu Jiang ⋅ Yuyu Zhang ⋅ Changqian Yu ⋅ Ping Nie ⋅ Wenhu Chen ⋅ Kelsey R Allen
Abstract
Aligning generative models with human preferences remains a central challenge. While RLHF is widely adopted, it relies on large-scale preference data. This limitation is particularly evident in evaluating instruction-guided image edits, even though humans can often learn such preferences from only a few examples. We present RewardClaw, an agentic framework that shifts preference alignment from training evaluation models on massive human preference datasets to test-time learning with only a small number of demonstrations. Instead of learning from large-scale annotations, RewardClaw aligns with human preferences by iteratively evolving a library of tools and skills from as few as 100 preference demonstrations. Given a source image, an edited image, and an editing instruction, an orchestrator agent selects the most relevant subset of tools and skills from the maintained library, and sub-agents invoke them to construct a reasoning chain that produces a reward score. By comparing predicted scores with ground-truth preferences and analyzing successes and failures in the reasoning process, the orchestrator automatically refines its library of tools and skills without additional human annotation. Using only 0.05\% of the EditReward preference data, RewardClaw achieves 47.4\% average accuracy on EditReward-Bench $K$=2/3/4 and GenAI-Bench, surpassing GPT-5 by 5.3 points. When used as a reward signal in FlowGRPO, RL-tuned models achieve 3.52 on ImgEdit-Bench.
Successful Page Load