标准强化学习人类反馈存在公平性缺陷,新方法提升模型对少数意见的响应能力。
Procedural Fairness Failures in RLHF from Preference Averaging
- 通过分离不同偏好模式优化,避免偏好平均导致的信号压制
- 在控制实验中对齐准确率从46.9%提升至67.9%,公平差距缩小
- 适合关注大模型公平性与决策偏见的研究者和开发者
强化学习从人类反馈(RLHF)将异质偏好聚合为单一奖励模型,隐含偏好同质性假设。当偏好存在差异时,该聚合方式引发程序性公平失效,即多数群体偏好主导奖励学习,而少数群体偏好被系统性低估。本文将对齐中的程序性公平定义为保留不同偏好信号,并指出标准RLHF因偏好平均而违反此原则。提出偏好感知的RLHF(PA-RLHF),在奖励学习阶段分离不同偏好模式的优化。在受控设置下,PA-RLHF将整体对齐准确率从46.9%提升至67.9%,最差与最佳对齐群体间的公平差距从15.9个百分点降至9.6个百分点。结果表明,即使在无噪声的控制环境中,奖励学习的结构设计也可引发程序性公平问题,对大型语言模型和自主系统具有直接影响,偏差奖励模型可能在序列决策中放大不平等。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural fairness failure where majority preference groups dominate reward learning while minority preferences are systematically under-represented. This work defines procedural fairness in alignment as preserving distinct preference signals during reward modeling and shows that standard RLHF violates this via preference averaging. Preference-Aware RLHF (PA-RLHF) is introduced, separating optimization across preference modes at the reward learning stage. In a controlled setting, PA-RLHF improves overall alignment accuracy from 46.9% to 67.9% and reduces the fairness gap between best and worst aligned groups from 15.9 to 9.6 percentage points. These results show that procedural fairness failures in alignment can arise from structural design choices in reward learning, even in controlled, noise-free settings, with direct implications for large language models and agentic systems, where biased reward models can compound inequities across sequential decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。