arXiv:2603.13891cs.CLcs.AI2026-03被引 1

19个大模型在文本标注中再现种族刻板印象,影响研究与决策。

Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation

  • 用姓名和方言线索测试模型标注偏差,发现系统性倾向。
  • 黑人姓名被评更攻击性(18/19模型),亚裔姓名显聪明但不自信(17/19)。
  • 适合关注算法偏见、数据治理与社会公平的研究者阅读。

大型语言模型(LLMs)正广泛用于学术研究、内容审核和招聘等场景的自动化文本标注。在涵盖19个模型、2项实验共超400万次标注判断下,我们发现文本中隐含的身份线索会系统性地导致标注结果偏向种族刻板印象。在39个标注任务中,含有黑人姓名的文本被18/19个模型评为更具攻击性,18/19个模型认为更爱八卦;亚裔姓名则呈现“玻璃天花板”特征:17/19模型认为更聪明,18/19模型认为更缺乏自信和社交性。阿拉伯姓名引发认知抬升与人际贬低,所有少数群体均被普遍评为自律性更低。在匹配方言实验中,同一句子以非裔美式英语书写时,被19个模型均判定为更不专业(均值差异-0.774)、更不具教育背景(-0.688)、更有毒(18/19)、更愤怒(19/19)。唯一例外是基于姓名的可雇佣性评分,微调后出现过度纠正,系统性偏好少数族裔申请人。这些发现表明,将大模型用作自动标注工具可能将社会性偏见直接嵌入支撑研究、治理与决策的数据集与度量中。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring. Across 19 LLMs and two experiments totaling more than 4 million annotation judgments, we show that subtle identity cues embedded in text systematically bias annotation outcomes in ways that mirror racial stereotypes. In a names-based experiment spanning 39 annotation tasks, texts containing names associated with Black individuals are rated as more aggressive by 18 of 19 models and more gossipy by 18 of 19. Asian names produce a bamboo-ceiling profile: 17 of 19 models rate individuals as more intelligent, while 18 of 19 rate them as less confident and less sociable. Arab names elicit cognitive elevation alongside interpersonal devaluation, and all four minority groups are consistently rated as less self-disciplined. In a matched dialect experiment, the same sentence is judged significantly less professional (all 19 models, mean gap $-0.774$), less indicative of an educated speaker ($-0.688$), more toxic (18/19), and more angry (19/19) when written in African American Vernacular English rather than Standard American English. A notable exception occurs for name-based hireability, where fine-tuning appears to overcorrect, systematically favoring minority-named applicants. These findings suggest that using LLMs as automated annotators can embed socially patterned biases directly into the datasets and measurements that increasingly underpin research, governance, and decision-making.

大模型偏见文本标注种族刻板数据伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。