arXiv:2604.22749cs.CL2026-04被引 1

LLM生成叙事中存在对全球多数族裔的刻板印象与边缘化,尤其在涉及美国身份提示时更严重。

Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities

论文配图:Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities
图 1 · 摘自论文原文
  • 通过开放文本生成测试不同国家身份的呈现方式
  • 少数族裔身份在中立故事中被严重低估,在从属角色中却超50倍高频出现
  • 即使替换为非美国籍提示,美国中心偏见仍持续存在,需重视全球视角

大型语言模型(LLMs)广泛应用于从日常到高风险的文本生成任务,包括模拟庇护申请者的访谈。尽管其应用潜力巨大,但存在将有害偏见编码并延续至全球非主导群体的风险。本文研究主流LLMs在开放式叙事生成提示下对民族身份的呈现。结果表明,存在持久的代表性伤害,包括刻板印象、身份抹除及单一维度刻画。少数族裔身份在中立故事中严重不足,在从属角色中则被过度呈现,其出现频率比主导角色高出逾50倍。当输入提示包含美国身份线索(如“American”)时,伤害程度显著加剧。值得注意的是,这些危害无法归因于迎合倾向,因为即便用非美国家身份替代美国线索,美国中心偏见依然存在。基于此,我们呼吁以全球多数族裔视角出发,开展文化伤害研究,警惕盲目采用以美国为中心的LLMs进行分类、监控与误述。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for text generation tasks from everyday use to high-stakes enterprise and government applications, including simulated interviews with asylum seekers. While many works highlight the new potential applications of LLMs, there are risks of LLMs encoding and perpetuating harmful biases about non-dominant communities across the globe. To better evaluate and mitigate such harms, more research examining how LLMs portray diverse individuals is needed. In this work, we study how national origin identities are portrayed by widely-adopted LLMs in response to open-ended narrative generation prompts. Our findings demonstrate the presence of persistent representational harms by national origin, including harmful stereotypes, erasure, and one-dimensional portrayals of Global Majority identities. Minoritized national identities are simultaneously underrepresented in power-neutral stories and overrepresented in subordinated character portrayals, which are over fifty times more likely to appear than dominant portrayals. The degree of harm is amplified when US nationality cues (e.g., ``American'') are present in input prompts. Notably, we find that the harms we identify cannot be explained away via sycophancy, as US-centric biases persist even when replacing US nationality cues with non-US national identities in the prompts. Based on our findings, we call for further exploration of cultural harms in LLMs through methodologies that center Global Majority perspectives and challenge the uncritical adoption of US-based LLMs for the classification, surveillance, and misrepresentation of the majority of our planet.

大模型偏见代表性伤害文化伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。