arXiv:2602.07799cs.LGcs.AI2026-02被引 1

让大模型奖励模型更公平,减少偏见生成。

Fairness Aware Reward Optimization

  • 在训练时加入公平性约束,确保奖励模型既准又公。
  • 实测显示偏见降低,且模型质量不降反升。
  • 适合关注大模型公平性的研究人员和开发者。

人类偏好数据中的群体偏差会通过奖励模型传递至对齐的大语言模型中,导致系统性不公平。我们提出公平感知奖励优化(Faro),一种在处理过程中施加人口均等、机会均等或反事实公平约束的框架。首次提供了大模型对齐中奖励层面公平性的理论分析,证明:(i) Faro 训练的奖励具备可调节松弛度的公平性证书;(ii) KL 正则微调引发的准确率-公平性权衡有形式化刻画,并证明公平性可从奖励传递到策略;(iii) 非空帕累托前沿存在。与预处理和后处理方法不同,Faro 确保奖励模型同时满足序数性(正确排序)、基数性(校准)和公平性。在多个大模型和基准测试中,Faro 显著降低了偏见和有害生成,同时保持或提升了模型性能。

原文摘要 · Abstract (English)

Demographic skews in human preference data propagate systematic unfairness through reward models into aligned LLMs. We introduce Fairness Aware Reward Optimization (Faro), an in-processing framework that trains reward models under demographic parity, equalized odds, or counterfactual fairness constraints. We provide the first theoretical analysis of reward-level fairness in LLM alignment, establishing: (i) provable fairness certificates for Faro-trained rewards with controllable slack; a (ii) formal characterization of the accuracy-fairness trade-off induced by KL-regularized fine-tuning, proving fairness transfers from reward to policy; and the (iii) existence of a non-empty Pareto frontier. Unlike pre- and post-processing methods, Faro ensures reward models are simultaneously ordinal (ranking correctly), cardinal (calibrated), and fair. Across multiple LLMs and benchmarks, Faro significantly reduces bias and harmful generations while maintaining or improving model quality.

公平性大模型奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。