测试大模型识别自闭症相关隐性歧视的能力,发现其易忽略语境中的伤害性
Beyond Keywords: Evaluating Large Language Model Classification of Nuanced Ableism
- 通过关键词匹配判断文本是否含歧视,忽略语境和语气
- 模型能识别自闭症相关词但常漏掉实际有害表达
- 适合关注AI偏见与伦理评估的研究者阅读
大型语言模型(LLMs)越来越多地应用于简历筛选、内容审核等决策任务,可能放大或压制特定观点。尽管已有研究指出LLMs存在残疾相关偏见,但对其如何理解或识别隐性能力主义仍知之甚少。本文评估了四种LLMs对针对自闭症群体的微妙能力主义内容的识别能力。研究对比了模型对相关术语的理解与其在具体语境中识别歧视性内容的效果。结果显示,模型虽能识别与自闭症相关的词汇,却常遗漏其中的伤害性或冒犯性含义。进一步的定性分析表明,模型依赖表面关键词匹配,导致语境误判;而人类标注者则会考虑语境、发言者身份及潜在影响。然而,模型与人类在标注框架上达成一致,表明二分类任务足以评估模型表现,与以往人类标注研究结果一致。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in decision-making tasks like résumé screening and content moderation, giving them the power to amplify or suppress certain perspectives. While previous research has identified disability-related biases in LLMs, little is known about how they conceptualize ableism or detect it in text. We evaluate the ability of four LLMs to identify nuanced ableism directed at autistic individuals. We examine the gap between their understanding of relevant terminology and their effectiveness in recognizing ableist content in context. Our results reveal that LLMs can identify autism-related language but often miss harmful or offensive connotations. Further, we conduct a qualitative comparison of human and LLM explanations. We find that LLMs tend to rely on surface-level keyword matching, leading to context misinterpretations, in contrast to human annotators who consider context, speaker identity, and potential impact. On the other hand, both LLMs and humans agree on the annotation scheme, suggesting that a binary classification is adequate for evaluating LLM performance, which is consistent with findings from prior studies involving human annotators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。