评测大模型奖励模型的群体公平性,发现主流模型普遍存在不公平问题。
Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models
- 用arXiv专家文本构建无提示一致性的公平性评测基准
- 所有测试奖励模型均存在统计显著的群体不公平性
- 性能越好的模型反而更公平,提示可优化公平性
随着大型语言模型(LLMs)日益强大且广泛可用,确保不同人口群体间的公平性(即群体公平性)成为关键的伦理问题。然而,当前关于LLM公平性和偏见的研究存在两大局限:一是传统机器学习中的群体公平性要求非敏感属性(如提示问题)在各群体间保持一致,但在实际场景中不同群体可能偏好不同提问方式,该要求不切实际;二是仅评估模型最终输出的公平性,未能定位偏见来源。事实上,偏见可能来自预训练、微调过程,以及强化学习人类反馈(RLHF)和学习到的奖励模型。因此,评估整个模型流水线中各组件的公平性有助于开发更好的缓解方法。针对上述问题,本文首次对学习到的奖励模型进行群体公平性基准评测。通过使用arXiv专家撰写文本,无需要求不同群体使用相同提示即可实现公平性评估。结果令人意外:所有被测奖励模型(如Nemotron-4-340B-Reward、ArmoRM-Llama3-8B-v0.1、GRM-llama3-8B-sftreg)均表现出统计显著的群体不公平性。此外,表现优异的模型(按标准性能指标衡量)往往展现出更好的群体公平性。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become increasingly powerful and accessible to human users, ensuring fairness across diverse demographic groups, i.e., group fairness, is a critical ethical concern. However, current fairness and bias research in LLMs is limited in two aspects. First, compared to traditional group fairness in machine learning classification, it requires that the non-sensitive attributes, in this case, the prompt questions, be the same across different groups. In many practical scenarios, different groups, however, may prefer different prompt questions and this requirement becomes impractical. Second, it evaluates group fairness only for the LLM's final output without identifying the source of possible bias. Namely, the bias in LLM's output can result from both the pretraining and the finetuning. For finetuning, the bias can result from both the RLHF procedure and the learned reward model. Arguably, evaluating the group fairness of each component in the LLM pipeline could help develop better methods to mitigate the possible bias. Recognizing those two limitations, this work benchmarks the group fairness of learned reward models. By using expert-written text from arXiv, we are able to benchmark the group fairness of reward models without requiring the same prompt questions across different demographic groups. Surprisingly, our results demonstrate that all the evaluated reward models (e.g., Nemotron-4-340B-Reward, ArmoRM-Llama3-8B-v0.1, and GRM-llama3-8B-sftreg) exhibit statistically significant group unfairness. We also observed that top-performing reward models (w.r.t. canonical performance metrics) tend to demonstrate better group fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。