arXiv:2502.12858cs.CLcs.AI2025-02NAACL被引 26

发现大模型奖励模型对非裔美国人语言存在系统性偏见

Rejected Dialects: Biases Against African American Language in Reward Models

  • 构建评估框架,对比白人主流英语与非裔美国人语言的模型偏好
  • 模型在非裔美国人语言上准确率低4%,更倾向排斥其表达
  • 揭示模型对话会主动转向主流英语,适合关注公平性的研究者

通过奖励模型实现偏好对齐有助于构建安全、有用且可靠的大型语言模型(LLMs)。然而,偏好判断中的主观性以及偏好数据收集中代表性不足可能引入新偏差,影响奖励模型的公平性。本文提出一种评估奖励模型中方言偏见的框架,并以非裔美国人语言(AAL)为例开展案例研究。通过对比白人主流英语(WME)与机器翻译及人工撰写的AAL语料,实验发现:当处理AAL文本时,模型平均准确率下降4%;更频繁地排斥符合AAL风格的文本;即使输入为AAL提示,也会引导对话转向WME。研究揭示了大模型发展中一个较少被关注阶段的表征性伤害,引发关于模型理想行为的伦理反思。

原文摘要 · Abstract (English)

Preference alignment via reward models helps build safe, helpful, and reliable large language models (LLMs). However, subjectivity in preference judgments and the lack of representative sampling in preference data collection can introduce new biases, hindering reward models' fairness and equity. In this work, we introduce a framework for evaluating dialect biases in reward models and conduct a case study on biases against African American Language (AAL) through several experiments comparing reward model preferences and behavior on paired White Mainstream English (WME) and both machine-translated and human-written AAL corpora. We show that reward models are less aligned with human preferences when processing AAL texts vs. WME ones (-4\% accuracy on average), frequently disprefer AAL-aligned texts vs. WME-aligned ones, and steer conversations toward WME, even when prompted with AAL texts. Our findings provide a targeted analysis of anti-AAL biases at a relatively understudied stage in LLM development, highlighting representational harms and ethical questions about the desired behavior of LLMs concerning AAL.

语言偏见公平性奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。