arXiv:2508.07111cs.CLcs.AI2025-08被引 4

通过置信度差异检测大模型在指代消解中的交叉性偏见。

Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

  • 构建新基准WinoIdentity,涵盖245,700个交叉身份提示
  • 发现模型对双重弱势群体的置信度低至40%差距
  • 揭示模型依赖记忆而非推理,存在价值对齐与有效性双重失效

大型语言模型(LLMs)在决策支持场景中广泛应用,但其可能反映并加剧社会偏见,引发身份相关伤害。现有研究多聚焦单一维度公平性评估,本文首次扩展至交叉性偏见分析。我们创建新基准WinoIdentity,基于WinoBias数据集,引入25个人口属性标记(包括年龄、国籍、种族等),与二元性别交叉,生成245,700个提示,用于评估50种不同偏见模式。聚焦因代表性不足导致的遗漏性伤害,从不确定性视角出发,提出核心指代置信度差异(Coreference Confidence Disparity)度量方法,衡量模型对不同交叉身份的置信程度差异。评估五款近期发布的模型发现,在身体类型、性取向、社会经济地位等维度上,置信度差异高达40%,尤其在反刻板印象情境中,对双重弱势身份群体最不确定。令人意外的是,主流或特权身份的置信度也下降,表明模型性能提升更源于记忆而非逻辑推理。这暴露了价值对齐与有效性的双重失败,可能共同导致社会危害。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved impressive performance, leading to their widespread adoption as decision-support tools in resource-constrained contexts like hiring and admissions. There is, however, scientific consensus that AI systems can reflect and exacerbate societal biases, raising concerns about identity-based harm when used in critical social contexts. Prior work has laid a solid foundation for assessing bias in LLMs by evaluating demographic disparities in different language reasoning tasks. In this work, we extend single-axis fairness evaluations to examine intersectional bias, recognizing that when multiple axes of discrimination intersect, they create distinct patterns of disadvantage. We create a new benchmark called WinoIdentity by augmenting the WinoBias dataset with 25 demographic markers across 10 attributes, including age, nationality, and race, intersected with binary gender, yielding 245,700 prompts to evaluate 50 distinct bias patterns. Focusing on harms of omission due to underrepresentation, we investigate bias through the lens of uncertainty and propose a group (un)fairness metric called Coreference Confidence Disparity which measures whether models are more or less confident for some intersectional identities than others. We evaluate five recently published LLMs and find confidence disparities as high as 40% along various demographic attributes including body type, sexual orientation and socio-economic status, with models being most uncertain about doubly-disadvantaged identities in anti-stereotypical settings. Surprisingly, coreference confidence decreases even for hegemonic or privileged markers, indicating that the recent impressive performance of LLMs is more likely due to memorization than logical reasoning. Notably, these are two independent failures in value alignment and validity that can compound to cause social harm.

交叉性偏见置信度分析指代消解大模型伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。