纠正推理模型训练中的优势估计偏差,提升性能
Your Group-Relative Advantage Is Biased
- 提出基于历史难度的自适应重加权机制
- 实验显示在5个数学推理任务上显著提升效果
- 适合研究强化学习后训练与推理优化的学者
基于验证器奖励的强化学习(RLVR)已成为大型语言模型推理能力后训练的常用方法,其中以GRPO及其变体为代表的分组方法广泛应用。这类方法依赖分组相对优势估计来避免学习价值函数,但其理论性质尚不清晰。本文首次揭示分组强化学习的根本问题:分组相对优势估计量相对于真实期望优势存在系统性偏差。理论上证明该偏差会低估难题的优势、高估易题的优势,导致探索与利用失衡。为此,我们提出历史感知的自适应难度加权(HA-DW),通过动态调整难度锚点和训练过程进行优势重加权。理论分析与五项数学推理基准上的实验证明,将HA-DW集成至GRPO及其变体中能持续提升性能。结果表明,修正优势估计偏差对鲁棒高效的RLVR训练至关重要。
原文摘要 · Abstract (English)
Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such as GRPO and its variants gaining broad adoption. These methods rely on group-relative advantage estimation to avoid learned critics, yet its theoretical properties remain poorly understood. In this work, we uncover a fundamental issue of group-based RL: the group-relative advantage estimator is inherently biased relative to the true (expected) advantage. We provide the first theoretical analysis showing that it systematically underestimates advantages for hard prompts and overestimates them for easy prompts, leading to imbalanced exploration and exploitation. To address this issue, we propose History-Aware Adaptive Difficulty Weighting (HA-DW), an adaptive reweighting scheme that adjusts advantage estimates based on an evolving difficulty anchor and training dynamics. Both theoretical analysis and experiments on five mathematical reasoning benchmarks demonstrate that HA-DW consistently improves performance when integrated into GRPO and its variants. Our results suggest that correcting biased advantage estimation is critical for robust and efficient RLVR training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。