arXiv:2603.15646cs.LGcs.AI2026-03被引 5

用分步优化多维度评分,让模型更精准地学习复杂任务。

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

  • 逐个优化评分维度,避免固定加权带来的偏差
  • 在4种规模模型上均提升性能与训练效率
  • 适合需要精细控制多个目标的生成任务

基于评分的强化学习(RLRR)通过结构化、多维度的上下文评分替代传统的标量偏好信号,扩展了人类反馈强化学习(RLHF)和可验证奖励(RLVR)。然而现有方法依赖固定权重将向量奖励线性压缩为标量,对评分设计敏感且无法捕捉各维度间相关性。本文提出交替式基于评分的强化学习(ARL-RR),通过一次优化一个语义评分元类别,无需固定标量聚合。理论上,奖励聚合会带来方差收缩效应,解释性能提升。进一步引入轻量级搜索机制,动态选择下一优化目标,使策略聚焦关键任务。在带有专家标注的HealthBench数据集上,ARL-RR在1.7B、4B、8B和14B不同规模模型上均优于标量方法,在模型性能与训练效率上表现一致更优。

原文摘要 · Abstract (English)

Reinforcement Learning with Rubric Rewards (RLRR) is a framework that extends conventional reinforcement learning from human feedback (RLHF) and verifiable rewards (RLVR) by replacing scalar preference signals with structured, multi-dimensional, contextual rubric-based evaluations. However, existing approaches in RLRR are limited to linearly compressing vector rewards into a scalar reward with a fixed weightings, which is sensitive to artificial score design and fails to capture correlations among reward dimensions. To overcome the limitations of reward aggregation, this work proposes Alternating Reinforcement Learning with Rubric Rewards (ARL-RR), a framework that eliminates the need for a fixed scalarization by optimizing one semantic rubric meta-class at a time. Theoretically, we show that reward aggregation induces a variance contraction effect, which helps explain the performance gains. We further introduce a lightweight, search-based adaptation procedure that selects the next meta-class dynamically based on task performance, enabling the policy to emphasize critical objectives and thereby improve the model performance. Empirically, our experiments on the HealthBench dataset with experts annotations demonstrate that ARL-RR uniformly outperforms scalarized methods in both model performance and training efficiency across different model scales (1.7B, 4B, 8B, and 14B).

强化学习多目标优化评分系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。