arXiv:2503.01623cs.HCcs.CL2025-03中稿 · CHI Conference on …被引 34

检测商业内容审核接口对群体仇恨言论的误判与漏判问题

Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations

  • 构建审计框架评估黑箱NLP系统的审核偏差
  • 五百万查询中发现所有接口均误判群体相关词汇
  • 适合平台方和政策制定者了解审核风险

商业内容审核API被宣传为应对网络仇恨言论的可扩展解决方案。然而,依赖这些API可能导致合法言论被过度压制(过审),或未能有效拦截有害言论(漏审)。本文提出一种审计黑箱NLP系统的方法,对五个主流商用内容审核API进行系统性评估。基于四个数据集的五百万次查询分析发现,各API频繁依赖“黑人”等群体身份词判断仇恨言论。尽管OpenAI和Amazon的服务表现稍优,但所有提供商均未能有效识别隐含式仇恨言论(如编码化表达),尤其针对LGBTQIA+群体。同时,它们普遍对反制言论、重新使用的贬义词及涉及黑人、LGBTQIA+、犹太人、穆斯林群体的内容实施过度审查。建议API提供方改进使用指引、阈值设置说明,并提高对其局限性的透明度。注意:本文包含仇恨言论术语,为保持透明度而复现。

原文摘要 · Abstract (English)

Commercial content moderation APIs are marketed as scalable solutions to combat online hate speech. However, the reliance on these APIs risks both silencing legitimate speech, called over-moderation, and failing to protect online platforms from harmful speech, known as under-moderation. To assess such risks, this paper introduces a framework for auditing black-box NLP systems. Using the framework, we systematically evaluate five widely used commercial content moderation APIs. Analyzing five million queries based on four datasets, we find that APIs frequently rely on group identity terms, such as ``black'', to predict hate speech. While OpenAI's and Amazon's services perform slightly better, all providers under-moderate implicit hate speech, which uses codified messages, especially against LGBTQIA+ individuals. Simultaneously, they over-moderate counter-speech, reclaimed slurs and content related to Black, LGBTQIA+, Jewish, and Muslim people. We recommend that API providers offer better guidance on API implementation and threshold setting and more transparency on their APIs' limitations. Warning: This paper contains offensive and hateful terms and concepts. We have chosen to reproduce these terms for reasons of transparency.

内容审核偏见检测黑箱审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。