arXiv:2508.06649cs.CL2025-08

检测大模型在性别、宗教等属性上的刻板与偏差,揭示生成内容的潜在风险。

Measuring Stereotype and Deviation Biases in Large Language Models

  • 通过生成人物档案,分析模型对不同群体的属性关联
  • 所有测试模型均在多个人群中表现出显著刻板与偏差
  • 适合关注AI伦理、公平性研究者参考

大型语言模型(LLMs)被广泛应用于多个领域,引发对其局限性和潜在风险的关注。本研究探究了两类模型可能存在的偏见:刻板印象偏见和偏离偏见。刻板印象偏见指模型持续将特定属性与某一社会群体关联;偏离偏见则体现为模型生成内容中的群体分布与真实世界分布之间的差异。我们让四款先进大模型生成个体人物档案,考察其对政治倾向、宗教信仰、性取向等属性与各人口群体的关联。实验结果表明,所有测试模型在多个群体中均表现出显著的刻板印象偏见和偏离偏见。研究揭示了模型在推断用户属性时产生的偏见,为理解大模型生成内容的潜在危害提供了依据。

原文摘要 · Abstract (English)

Large language models (LLMs) are widely applied across diverse domains, raising concerns about their limitations and potential risks. In this study, we investigate two types of bias that LLMs may display: stereotype bias and deviation bias. Stereotype bias refers to when LLMs consistently associate specific traits with a particular demographic group. Deviation bias reflects the disparity between the demographic distributions extracted from LLM-generated content and real-world demographic distributions. By asking four advanced LLMs to generate profiles of individuals, we examine the associations between each demographic group and attributes such as political affiliation, religion, and sexual orientation. Our experimental results show that all examined LLMs exhibit both significant stereotype bias and deviation bias towards multiple groups. Our findings uncover the biases that occur when LLMs infer user attributes and shed light on the potential harms of LLM-generated outputs.

大模型偏见刻板印象伦理风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。