首个平衡名人度与人口统计的引用归属测试集,揭示大模型存在系统性偏见。
Attribution Bias in Large Language Models
- 构建平衡名人度和人口统计特征的引用归属数据集AttriBench。
- 11个主流大模型在种族、性别上归属准确率差异显著,部分群体归因率低至40%。
- 发现模型会完全省略引用(抑制现象),且该问题在不同群体间分布不均。
随着大语言模型(LLMs)越来越多用于搜索和信息检索,准确归因内容来源变得至关重要。本文提出AttriBench,首个在名人度和人口统计特征上均衡的引文归属基准数据集。通过显式平衡作者知名度与人口统计属性,AttriBench使我们能够对引文归属中的群体偏见进行受控研究。我们在不同提示设置下评估了11个广泛使用的LLMs,发现即使在前沿模型中,引文归属仍是挑战性任务。观察到种族、性别及交叉群体间归属准确率存在显著系统性差异。我们进一步引入并分析‘抑制’这一独特失效模式——即模型在可获取作者信息时仍完全省略归属。结果显示该现象普遍存在且在不同群体间分布不均,揭示了传统准确率指标无法捕捉的系统性偏见。本研究将引文归属定位为衡量大模型表征公平性的基准。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are increasingly used to support search and information retrieval, it is critical that they accurately attribute content to its original authors. In this work, we introduce AttriBench, the first fame- and demographically-balanced quote attribution benchmark dataset. Through explicitly balancing author fame and demographics, AttriBench enables controlled investigation of demographic bias in quote attribution. Using this dataset, we evaluate 11 widely used LLMs across different prompt settings and find that quote attribution remains a challenging task even for frontier models. We observe large and systematic disparities in attribution accuracy between race, gender, and intersectional groups. We further introduce and investigate suppression, a distinct failure mode in which models omit attribution entirely, even when the model has access to authorship information. We find that suppression is widespread and unevenly distributed across demographic groups, revealing systematic biases not captured by standard accuracy metrics. Our results position quote attribution as a benchmark for representational fairness in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。