通过隐式偏好信号提升GRPO在复杂推理中的表现
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
- 从组内奖励排名构建对比正则,无需额外标注
- 显著抑制低质量推理路径,减少长度偏差
- 适合需要精细推理对齐的大型语言模型训练
强化学习已成为对齐大语言模型在复杂推理任务上的主要范式,其中群体相对策略优化(GRPO)广泛应用于大规模后训练。然而,GRPO在重推理场景中存在结构性缺陷:序列级优势归一化引入系统性长度偏差,低质量轨迹的惩罚被稀释,标量目标也丢弃了组内奖励排序所蕴含的丰富成对偏好信息。因此,来自高成本采样的宝贵监督信号未能充分使用。我们提出AMIR-GRPO,通过直接从组内奖励排序构建类似DPO的隐式对比正则,无需额外标注。该机制增强对低奖励轨迹的抑制,减轻响应级长度偏差,并将每个采样组转化为更密集的监督约束集。在多个数学推理基准上,AMIR-GRPO持续优于强基线GRPO,使正确与错误推理链区分更清晰,并在标准GRPO无法解决的子集上实现更广覆盖提升。
原文摘要 · Abstract (English)
Reinforcement learning has become the primary paradigm for aligning large language models (LLMs) on complex reasoning tasks, with group relative policy optimization (GRPO) widely used in large-scale post-training. However, GRPO faces structural limitations in reasoning-heavy settings: sequence-level advantage normalization introduces systematic length bias, penalties for low-quality trajectories are diluted, and the scalar objective discards rich pairwise preference information embedded in within-group reward rankings. As a result, valuable supervision from costly rollouts remains underutilized. We propose AMIR-GRPO, which augments GRPO with an implicit DPO-style contrastive regularizer constructed directly from intra-group reward rankings, requiring no additional annotations. This mechanism amplifies suppression of low-reward trajectories, attenuates response-level length bias, and transforms each rollout group into a denser set of supervision constraints. Across multiple mathematical reasoning benchmarks, AMIR-GRPO consistently outperforms strong GRPO baselines, yields clearer separation between correct and incorrect reasoning chains, and delivers broader coverage gains beyond the subset of instances solved by standard GRPO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。