Skip to yearly menu bar Skip to main content


Poster Thu, Oct 8, 2026 • 11:00 AM – 1:00 PM PDT Grand Ballroom #118

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards

Mengjie Ren ⋅ Jie Lou ⋅ Boxi Cao ⋅ Xueru Wen ⋅ Hongyu Lin ⋅ Xianpei Han ⋅ Le Sun ⋅ XingYu Li ⋅ Yaojie Lu

Abstract

Log in and register to view live content