arXiv:2509.03672cs.LGstat.ML2025-09AAAI被引 1

提出共享特征框架,更好兼顾少数群体偏好,提升公平性与性能。

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

  • 通过学习跨群体共享特征,替代分组建模,提升泛化能力。
  • 在多任务实验中,胜率最高提升20%,尤其改善少数群体表现。
  • 适合关注模型公平性、需处理多元偏好的研究者使用。

统一奖励的强化学习从人类反馈(RLHF)通过单一奖励模型表示所有标注者的偏好,无法捕捉子群体间的意见多样性,无意中偏向主流群体。当前最优方法MaxMin-RLHF通过学习分组奖励模型并优化最低得分群体来提升公平性,但其在最低奖励群体为少数时性能显著下降。为此,我们提出新框架SharedRep-RLHF,核心在于学习不同群体间标注中的共享特征,而非分别建模各组。我们证明MaxMin-RLHF在学习共享特征上是严格次优的,并量化了SharedRep-RLHF的样本复杂度。在多种自然语言任务上的实验表明,相比MaxMin-RLHF,SharedRep-RLHF在胜率上最高提升20%。

原文摘要 · Abstract (English)

Uniform-reward reinforcement learning from human feedback (RLHF), which trains a single reward model to represent the preferences of all annotators, fails to capture the diversity of opinions across sub-populations, inadvertently favoring dominant groups. The state-of-the-art, MaxMin-RLHF, addresses this by learning group-specific reward models, and by optimizing for the group receiving the minimum reward, thereby promoting fairness. However, we identify that a key limitation of MaxMin-RLHF is its poor performance when the minimum-reward group is a minority. To mitigate this drawback, we introduce a novel framework, termed {\em SharedRep-RLHF}. At its core, SharedRep-RLHF learns and leverages {\em shared traits} in annotations among various groups, in contrast to learning separate reward models across groups. We first show that MaxMin-RLHF is provably suboptimal in learning shared traits, and then quantify the sample complexity of SharedRep-RLHF. Experiments across diverse natural language tasks showcase the effectiveness of SharedRep-RLHF compared to MaxMin-RLHF with a gain of up to 20% in win rate.

RLHF公平性共享表征多群体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。