arXiv:2605.24384cs.CLcs.AI2026-05被引 1

对比评测放大语言模型对非裔英语的隐性偏见,暴露现有评估的局限性。

Side-by-side Comparison Amplifies Dialect Bias in Language Models

论文配图:Side-by-side Comparison Amplifies Dialect Bias in Language Models
图 1 · 摘自论文原文
  • 通过对比标准英语与非裔英语推文,发现模型偏见在对比场景下加剧。
  • 显式标注方言后偏见进一步加重,且部分缓解方法在对比场景中失效。
  • 提醒需改进评估框架,尤其关注对比决策场景中的公平性问题。

语言模型(LMs)可能在无方言标签的情况下表现出基于方言的隐性偏见,称为隐性方言偏见。本文通过评估意图相同的标准美式英语(SAE)和非裔美国人口语英语(AAVE)推文,量化在线话语中的隐性方言偏见,使用社会心理学研究中关于种族偏见的刻板印象特征作为指标。已有研究表明,当单独评估时,模型更倾向于将负面刻板印象关联到AAVE;而令人意外的是,在并列对比的场景下,这种偏见显著加剧,该场景更贴近模型用于候选人排序等高影响决策的实际应用。当显式标注方言时,偏见进一步恶化。尽管商业开发者投入大量努力减轻偏见,但此现象仍普遍存在。我们发现反事实公平微调可缓解部分刻板印象的偏见,但在并列对比下效果不一致。结果表明,现有隐性偏见评估设置可能低估其严重性,尤其是在对比情境中。此外,即使经过安全对齐微调,显性方言偏见依然明显,说明该问题仍未解决,亟需更鲁棒的评估与缓解框架。

原文摘要 · Abstract (English)

Language models (LMs) can exhibit biases based on variations in their dialects, even in the absence of a dialect label, a behavior known as covert dialect bias. In this work, we quantify covert dialect bias in online discourse by evaluating how LMs associate stereotypical traits (derived from social psychology research on racial bias) with intent-equivalent tweets in Standard American English (SAE) and African-American Vernacular English (AAVE). While prior work shows that LMs associate more negative stereotypes with AAVE when evaluating tweets in isolation, we are surprised to find that this bias is significantly exacerbated when SAE / AAVE tweet pairs are compared side by side, a setting that more closely reflects high-impact decision making contexts in which models are used to rank candidates. The bias only worsens when dialect labels are explicitly specified. This is striking, given the extensive efforts from commercial developers to mitigate bias in their LMs. Encouragingly, we show that counterfactual fairness finetuning can mitigate covert dialect bias for some stereotypical traits, reducing average disparities when evaluating tweets in isolation, however, these improvements do not consistently hold across traits when evaluating SAE / AAVE tweets side by side. Our findings show that existing evaluation settings for covert dialect bias may underestimate its severity, specifically in contrastive settings. Additionally, overt dialect bias remains pronounced even after safety aligned finetuning, indicating that it remains an unresolved problem, and motivates the need for more robust evaluation and mitigation frameworks.

语言模型偏见检测方言公平性对比评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。