揭露AI毒性检测模型对非裔英语的系统性偏见,用互动工具揭示算法歧视如何被政策放大。
How AI Fails: An Interactive Pedagogical Tool for Demonstrating Dialectal Bias in Automated Toxicity Models
- 对比非裔英语与标准英语文本的毒性评分差异
- 非裔英语文本平均毒性分高出1.8倍,身份仇恨分高8.8倍
- 通过可调阈值工具让公众直观感受算法偏见的现实危害
随着AI内容审核在日常生活中日益普及,人们常调侃称“AI有偏见”。这一玩笑背后隐藏着深层忧虑:一条被标记为“不当”的帖子,是否只是算法偏见的受害者?本文采用双重方法研究此问题。首先,对广泛使用的毒性模型(unitary/toxic-bert)进行量化基准测试,评估其在非裔英语(AAE)与标准美式英语(SAE)文本上的表现差异。结果显示,该模型对AAE文本的毒性评分平均高出1.8倍,身份仇恨评分高达8.8倍。其次,本文提出一个交互式教育工具,核心机制是用户可调节的“敏感度阈值”,揭示偏见不仅存在于评分本身,更在于人类设定的看似中立的政策最终导致歧视性后果。本研究既提供歧视性影响的统计证据,也推出面向公众的工具,旨在提升批判性AI素养。
原文摘要 · Abstract (English)
Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that "the AI is biased". While this is often said jokingly, the light-hearted remark reflects a deeper concern. How can we be certain that an online post flagged as "inappropriate" was not simply the victim of a biased algorithm? This paper investigates this problem using a dual approach. First, I conduct a quantitative benchmark of a widely used toxicity model (unitary/toxic-bert) to measure performance disparity between text in African-American English (AAE) and Standard American English (SAE). The benchmark reveals a clear, systematic bias: on average, the model scores AAE text as 1.8 times more toxic and 8.8 times higher for "identity hate". Second, I introduce an interactive pedagogical tool that makes these abstract biases tangible. The tool's core mechanic, a user-controlled "sensitivity threshold," demonstrates that the biased score itself is not the only harm; instead, the more-concerning harm is the human-set, seemingly neutral policy that ultimately operationalises discrimination. This work provides both statistical evidence of disparate impact and a public-facing tool designed to foster critical AI literacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。