AI内容审核难辨歧视语是否被社群重新定义,易误伤边缘群体表达。
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language

- 通过标注社区内歧视语使用数据,分析用户对重用词的判断差异
- 跨群体内部一致度低,同一社群成员也常对是否构成仇恨言论分歧
- 现有AI系统与用户判断偏差大,适合需理解语境的伦理与社会研究者
许多边缘化社群在线上重用歧视性词汇,作为团结、身份认同与共同经历的表达。然而,当前基于AI的内容审核工具难以区分此类重用与仇恨言论,导致边缘群体声音被压制。本研究通过定量与定性方法,考察了LGBTQIA+、黑人及女性社群对特定歧视词(如f-word、n-word、b-word)重用的态度。我们收集并分析了一个经标注的在线歧视语使用语料库,包含标注者对含歧视词文本是否应标记为仇恨言论的判断,以及使用语境特征。在所有社群和标注问题中,标注者间一致性较低,表明同社群内部也存在显著分歧。缺乏明确的身份与意图信号时,甚至同群成员也无法达成共识。半结构化访谈显示,个人生活经历差异是造成这种变异的原因之一。结果表明,标注者判断与Perspective API生成的自动化仇恨言论评估严重不匹配。此外,文本中词汇是否具有贬义、是否自指等特征更影响是否报告为仇恨言论。这些发现凸显了边缘社群对歧视语解读的高度主观性与语境依赖性。
原文摘要 · Abstract (English)
Reclaimed slur usage is a common and meaningful practice online for many marginalized communities. It serves as a source of solidarity, identity, and shared experience. However, contemporary automated and AI-based moderation tools for online content largely fail to distinguish between reclaimed and hateful uses of slurs, resulting in the suppression of marginalized voices. In this work, we use quantitative and qualitative methods to examine the attitudes of social media users in LGBTQIA+, Black, and women communities around reclaimed slurs targeting our focus groups including the f-word, n-word, and b-word. With social media users from these communities, we collect and analyze an annotated online slur usage corpus. The corpus includes annotators' perceptions of whether an online text containing a slur should be flagged as hate speech, as well as contextual features of the slur usage. Across all communities and annotation questions, we observe low inter-annotator agreement, indicating substantial disagreement among in-group annotators. This is compounded by the fact that, absent clear contextual signals of identity and intent, even in-group members may disagree on how to interpret reclaimed slur usage online. Semi-structured interviews with annotators suggest that differences in lived experience and personal history contribute to this variation as well. We find poor alignment between annotator judgments and automated hate speech assessments produced by Perspective API. We further observe that certain features of a text such as whether the slur usage was derogatory and if the slur was targeted at oneself are more associated with whether annotators report the text as hate speech. Together, these findings highlight the inherent subjectivity and contextual nature of how marginalized communities interpret slurs online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。