AI 日报hiw3c.com

Multi-Reward RL, Part 3: GDPO + CISPO + REPO-R at 27B, and the Advantage Floor That Stopped Learning

BestBlogs·AI 高分精选 www.bestblogs.dev 网页快照

Multi-Reward RL, Part 3: GDPO + CISPO + REPO-R at 27B, and the Advantage Floor That Stopped Learning

This study explores how different reward aggregation and policy optimization methods (GDPO, CISPO, REPO-R) affect the performance of a 27B model on a structured answer task, revealing that while training curves may look similar, the underlying learning dynamics and generalization capabilities vary significantly, with a specific 'advantage floor' guard causing a major failure in a harder environment.