arXiv:2603.10535cs.LGcs.CL2026-03被引 4

用乘法重标度解决大模型推理冗余问题,不损失性能。

Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning

  • 提出乘法式奖励重标度机制,动态控制输出长度。
  • 在多种强化学习场景下,显著减少冗长输出且保持性能不变。
  • 适合追求简洁高效推理的模型训练者使用。

强化学习能显著提升大语言模型能力,但面临严重问题:长度膨胀,即模型为最大化奖励而变得冗长或低效推理。现有方法难以通用且无损地解决此问题,因加性惩罚会引入补偿效应形成优化捷径,而启发式门控策略泛化性差,仅适用于二元反馈。为此,本文提出分组相对奖励重标度(GR$^3$),将长度控制重构为乘法重标度范式,建立通用、连续且依赖奖励的门控机制。为确保无损优化,引入分组相对正则化与优势感知校准,动态调整实例难度下的长度预算,并保留高质量轨迹的优势信号。实验证明,在RLHF与RLVR设置中,GR$^3$维持与标准GRPO相当的训练动态与下游性能,显著缓解长度膨胀,优于当前最优的长度正则化基线。

原文摘要 · Abstract (English)

Reinforcement learning significantly enhances LLM capabilities but suffers from a critical issue: length inflation, where models adopt verbosity or inefficient reasoning to maximize rewards. Prior approaches struggle to address this challenge in a general and lossless manner, primarily because additive penalties introduce a compensatory effect that creates optimization shortcuts, while heuristic gating strategies lack generality beyond binary feedback. To bridge this gap, we present Group Relative Reward Rescaling (GR$^3$), which reframes length control as a multiplicative rescaling paradigm, effectively establishing a generalized, continuous, and reward-dependent gating mechanism. To further ensure lossless optimization, we incorporate group-relative regularization and advantage-aware calibration, which dynamically adapt length budgets to instance difficulty and preserve the advantage signal of high-quality trajectories. Empirically, across both RLHF and RLVR settings, GR$^3$~maintains training dynamics and downstream performance comparable to standard GRPO while significantly mitigating length inflation, outperforming state-of-the-art length-regularized baselines.

强化学习长度控制大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。