arXiv:2505.23349cs.LGcs.AI2025-05ACL被引 12

从资源分配视角提升强化学习中的奖励公平性,避免偏见影响模型对齐。

Towards Reward Fairness in RLHF: From a Resource Allocation Perspective

  • 将偏好学习视为资源分配问题,用公平性约束优化奖励分配。
  • 在验证和强化学习场景中均实现更公平的奖励模型与策略模型。
  • 不针对特定偏见设计,通用性强,适合追求公平对齐的LLM研究者。

奖励作为人类偏好的代理,在基于人类反馈的强化学习(RLHF)中起关键作用。然而,若奖励本身存在固有偏差,将对大语言模型(LLMs)的对齐产生负面影响。本文将各类奖励偏差统一定义为奖励不公平问题。提出一种无偏见依赖的方法,从资源分配视角解决奖励公平性问题,无需针对每种偏差专门设计,却能有效缓解多种偏差。具体地,将偏好学习建模为资源分配问题,将奖励视为需分配的资源,同时权衡分配的效用与公平性。提出公平性正则化与公平性系数两种方法,分别用于构建公平奖励模型和公平策略模型。在验证与强化学习场景中的实验表明,该方法使大语言模型更公平地对齐人类偏好。

原文摘要 · Abstract (English)

Rewards serve as proxies for human preferences and play a crucial role in Reinforcement Learning from Human Feedback (RLHF). However, if these rewards are inherently imperfect, exhibiting various biases, they can adversely affect the alignment of large language models (LLMs). In this paper, we collectively define the various biases present in rewards as the problem of reward unfairness. We propose a bias-agnostic method to address the issue of reward fairness from a resource allocation perspective, without specifically designing for each type of bias, yet effectively mitigating them. Specifically, we model preference learning as a resource allocation problem, treating rewards as resources to be allocated while considering the trade-off between utility and fairness in their distribution. We propose two methods, Fairness Regularization and Fairness Coefficient, to achieve fairness in rewards. We apply our methods in both verification and reinforcement learning scenarios to obtain a fairness reward model and a policy model, respectively. Experiments conducted in these scenarios demonstrate that our approach aligns LLMs with human preferences in a more fair manner.

奖励公平性强化学习大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。