解决大模型生成中多维度评分失衡问题,提升综合质量
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

- 通过逆向奖励投影估算各评分维度饱和度,动态调整奖励方向
- 在6个基准上18次对比均超越静态聚合基线,显著提升平衡性
- 适合需要多维质量控制的生成模型优化,如对话、写作场景
大语言模型的开放式生成通常需要多维度评分标准来全面评估质量并指导强化学习优化。然而,该训练范式存在一个关键困境:不同评分维度间奖励分布严重失衡。即使模型训练后整体奖励较高,仍可能在某些维度表现严重不足,直接降低用户体验。为此,我们提出Focal Reward,一种新型目标函数,用于在基于评分标准的强化学习中实现自动平衡。具体而言,首先利用逆向奖励投影机制估计评分维度的饱和程度,作为校准奖励方向的基础;随后设计带有自动重加权系数的目标函数,实现细粒度平衡。在三个模型规模和六个基准上的大量实验表明,我们的Focal Reward方法在全部18次模型-基准组合中均优于最强的静态聚合基线。滚动分析、机制分析和消融实验进一步证明,性能提升源于对仍有改进空间维度的在线、饱和感知资源再分配。
原文摘要 · Abstract (English)
The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical dilemma inherent in this training paradigm is the imbalanced reward polarization along different rubric dimensions. Under this bottleneck, even if LLMs achieve relatively high rewards after training, they may still exhibit severe deficiencies in certain dimensions, leading to a direct deterioration in user experience. To address this problem, we propose Focal Reward, a novel objective to automatically balance the training of reinforcement learning under rubric-based rewards. Specifically, we first leverage an inverse reward projection mechanism to estimate the saturation degree of each criterion in the rubric, which forms the basis to calibrate the reward direction. Then, the final objective is designed with an automatically reweighting coefficient for each criterion to achieve the fine-grained balancing. Extensive experiments across three model scales and six benchmarks demonstrate that our Focal Reward method outperforms the strongest static aggregation baseline in all 18 model-benchmark comparisons. Rollout, mechanism, and ablation analyses further show that these gains arise from online, saturation-aware reallocation toward rubrics that still have room for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。