首个聚焦自闭症歧视语言的语境检测数据集,填补NLP领域空白。
AUTALIC: A Dataset for Anti-AUTistic Ableist Language In Context
- 构建2400条自闭症相关句子,含上下文与神经多样性专家标注
- 现有大模型在识别此类语言时表现不佳,难以匹配人类判断
- 适合研究歧视语言、神经多样性及标注分歧的研究者使用
随着对自闭症和能力主义理解的加深,针对自闭症人士的歧视性语言也日益受到关注。这类语言在自然语言处理中极具挑战性,因其微妙且依赖语境。然而,反自闭症能力主义语言的检测仍处于探索阶段,现有工具常无法捕捉其细微表达。本文提出AUTALIC,首个专注于反自闭症能力主义语言语境检测的基准数据集,包含从Reddit收集的2400条自闭症相关句子及其上下文,由具备神经多样性背景的训练专家进行标注。全面评估表明,当前语言模型(包括最先进的大模型)在识别此类语言时表现有限,难以与人类判断对齐,凸显其在该领域的不足。我们公开发布AUTALIC及个体标注结果,为研究能力主义、神经多样性及标注分歧提供宝贵资源。该数据集是构建更具包容性与语境感知的NLP系统的关键一步。
原文摘要 · Abstract (English)
As our understanding of autism and ableism continues to increase, so does our understanding of ableist language towards autistic people. Such language poses a significant challenge in NLP research due to its subtle and context-dependent nature. Yet, detecting anti-autistic ableist language remains underexplored, with existing NLP tools often failing to capture its nuanced expressions. We present AUTALIC, the first benchmark dataset dedicated to the detection of anti-autistic ableist language in context, addressing a significant gap in the field. The dataset comprises 2,400 autism-related sentences collected from Reddit, accompanied by surrounding context, and is annotated by trained experts with backgrounds in neurodiversity. Our comprehensive evaluation reveals that current language models, including state-of-the-art LLMs, struggle to reliably identify anti-autistic ableism and align with human judgments, underscoring their limitations in this domain. We publicly release AUTALIC along with the individual annotations which serve as a valuable resource to researchers working on ableism, neurodiversity, and also studying disagreements in annotation tasks. This dataset serves as a crucial step towards developing more inclusive and context-aware NLP systems that better reflect diverse perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。