AI fails to detect disability hate as well as disabled people do,且解释常出错
"Cold, Calculated, and Condescending": How AI Identifies and Explains Ableism Compared to Disabled People
- 用200条针对残障人士的社交媒体评论,对比AI与残障者对歧视内容的判断
- AI低估毒性程度,对残障歧视的识别不一致,解释缺乏细节且常武断
- 呼吁在算法审核中纳入残障者多元视角,避免技术加剧偏见
残障人士在线上常遭遇歧视性言论与微侵犯。当前平台主要依赖机器学习模型进行内容审核,但关于这些模型识别残障歧视的有效性及其判断是否与残障者一致,尚不清楚。为此,我们构建了首个包含200条针对残障人士的社交媒体评论的数据集,并让前沿AI模型(如毒性分类器、大语言模型)对每条评论评分并解释其判断依据。同时,我们招募190名参与者进行类似评分与解释,并评估大语言模型的解释质量。混合方法分析揭示了显著差距:相比残障者评分,AI低估了言论的毒性,其对残障歧视的识别零散且不一致。尽管大语言模型能识别部分偏见,但其解释存在不足——缺乏细致分析、做出错误假设,且语气显得评判而非教育。本文讨论了未来设计残障歧视审核系统的挑战与机遇,强调必须将交叉性残障视角纳入AI开发过程。
原文摘要 · Abstract (English)
People with disabilities (PwD) regularly encounter ableist hate and microaggressions online. These spaces are generally moderated by machine learning models, but little is known about how effectively AI models identify ableist speech and how well their judgments align with PwD. To investigate this, we curated a first-of-its-kind dataset of 200 social media comments targeted towards PwD, and prompted state-of-the art AI models (i.e., Toxicity Classifiers, LLMs) to score toxicity and ableism for each comment, and explain their reasoning. Then, we recruited 190 participants to similarly rate and explain the harm, and evaluate LLM explanations. Our mixed-methods analysis highlighted a major disconnect: AI underestimated toxicity compared to PwD ratings, while its ableism assessments were sporadic and varied. Although LLMs identified some biases, its explanations were flawed--they lacked nuance, made incorrect assumptions, and appeared judgmental instead of educational. Going forward, we discuss challenges and opportunities in designing moderation systems for ableism, and advocate for the involvement of intersectional disabled perspectives in AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。