arXiv:2510.21011cs.HCcs.AI2025-10被引 3

4大模型生成职业人物画像存在显著种族性别偏见,真实数据对比揭示系统性偏差。

Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

  • 用4大模型生成41个职业150万个人物,对比美国劳工统计局数据
  • 白人少31个百分点,黑人少9个百分点,西裔多17个百分点,亚裔多12个百分点
  • 跨模型重复出现偏差,说明是系统性结构问题而非单个模型缺陷

随着生成式AI被广泛用于描绘职业角色,理解其在种族与性别表征上的偏差至关重要。我们对GPT-4、Gemini 2.5、DeepSeek V3.1和Mistral-medium四款主流大语言模型,在41个美国职业中生成的超过150万条职业人物画像进行了审计。将这些生成结果与美国劳工统计局(BLS)数据对比发现,模型生成的人口统计特征变异程度低于真实数据,实质上将每个职业压缩至主导人群画像,未能体现真实人口多样性。通过偏差分解发现:白人(-31百分点)、黑人(-9百分点)普遍被低估,而西裔(+17百分点)、亚裔(+12百分点)被高估,且刻板印象加剧了既有职业隔离现象。偏差极端案例包括:家政人员几乎全被呈现为西裔,许多职业中黑人几乎完全消失。这些模式在具有不同文化背景的模型间反复出现,表明偏差源于共享的结构性根源,而非模型特异性产物。我们主张,评估生成式AI需建立框架,以审视合成人口如何系统性重塑社会角色中的群体可见性。

原文摘要 · Abstract (English)

As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational biases is critical. We audit over 1.5 million occupational personas generated by four major large language models (GPT-4, Gemini 2.5, DeepSeek V3.1, and Mistral-medium) across 41 U.S. occupations. Comparing these personas against U.S. Bureau of Labor Statistics (BLS) data, we find that models generate demographics with less variation than real-world data, functionally compressing each occupation toward a dominant demographic profile rather than representing population-level variation. A shift/exaggeration decomposition reveals the structure of these distortions: White (-31 percentage points) and Black (-9 pp) workers are consistently underrepresented, while Hispanic (+17 pp) and Asian (+12 pp) workers are overrepresented, with stereotype exaggeration amplifying existing occupational segregation. These distortions are often extreme, including near-total portrayals of housekeepers as Hispanic and the near-erasure of Black workers from many occupations. Because these patterns recur across models with different institutional and cultural origins, they suggest shared structural sources of bias rather than model-specific artifacts. We argue that auditing generative AI requires evaluation frameworks that examine how synthetic populations systematically reshape demographic visibility across social roles.

AI偏见职业画像生成模型种族差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。