arXiv:2505.15553cs.CLcs.AI2025-05被引 7

主流问答数据集存在性别、宗教和地域偏见,可能让大模型学坏。

Social Bias in Popular Question-Answering Benchmarks

  • 分析30篇论文与20个数据集,发现创建者信息不透明
  • 多数数据集含性别、宗教、地理偏见,仅一个(WinoGrande)主动应对
  • 揭示评估基准本身可能催生模型偏见,适合关注AI公平性的研究者

问答(QA)与阅读理解(RC)基准常用于评估大语言模型的知识获取与复现能力。然而,我们证明主流的QA与RC基准在覆盖不同人口群体或地区方面缺乏代表性。通过对30篇基准论文进行内容分析,以及对20个相应数据集的量化分析,我们考察了(1)基准创建中涉及的人员,(2)基准是否表现出社会偏见,或是否有措施应对与防范,(3)创建者与标注者的身份是否与内容中的偏见相关。大多数被分析的基准论文未提供关于标注者等参与人员的充分信息。值得注意的是,仅有一个(WinoGrande)明确报告了针对社会代表性问题所采取的措施。数据集分析显示,广泛存在的百科类、常识性及学术性基准中存在性别、宗教和地理偏见。本研究为日益增长的对人工智能评估实践的批评增添了新证据,揭示有偏基准可能成为大语言模型偏见的来源,通过激励有偏推理策略产生影响。

原文摘要 · Abstract (English)

Question-answering (QA) and reading comprehension (RC) benchmarks are commonly used for assessing the capabilities of large language models (LLMs) to retrieve and reproduce knowledge. However, we demonstrate that popular QA and RC benchmarks do not cover questions about different demographics or regions in a representative way. We perform a content analysis of 30 benchmark papers and a quantitative analysis of 20 respective benchmark datasets to learn (1) who is involved in the benchmark creation, (2) whether the benchmarks exhibit social bias, or whether this is addressed or prevented, and (3) whether the demographics of the creators and annotators correspond to particular biases in the content. Most benchmark papers analyzed provide insufficient information about those involved in benchmark creation, particularly the annotators. Notably, just one (WinoGrande) explicitly reports measures taken to address social representation issues. Moreover, the data analysis revealed gender, religion, and geographic biases across a wide range of encyclopedic, commonsense, and scholarly benchmarks. Our work adds to the mounting criticism of AI evaluation practices and shines a light on biased benchmarks being a potential source of LLM bias by incentivizing biased inference heuristics.

AI伦理评估基准偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。