审计Twitch自动审核系统,发现其对仇恨言论识别率低且误杀正常表达。
Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
- 用真实直播场景测试10万+评论,验证自动审核效果
- 94%仇恨内容未被识别,依赖辱骂词触发才生效
- 89.5%正常用语被误删,尤其在教育或赋权语境中
为应对内容审核需求,线上平台采用自动化系统。以Twitch为代表的实时互动平台对审核系统的延迟提出更高要求。尽管此类系统广泛应用,但其有效性仍缺乏了解。本文通过创建直播账号作为测试环境,利用Twitch API发送来自4个数据集的超10.7万条评论,审计其自动审核工具AutoMod对仇恨言论的识别能力。实验显示,高达94%的明显包含性别歧视、种族主义、能力歧视和恐同言论的内容绕过审核;而仅在添加辱骂词后,系统才实现100%删除,表明其严重依赖关键词作为判断信号。同时,违反社区准则的是,系统将高达89.5%的使用敏感词汇的非恶意内容(如教育性或赋权性表达)错误屏蔽。结果揭示AutoMod存在显著能力缺口,强调理解上下文对审核系统至关重要。
原文摘要 · Abstract (English)
To meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement($\textit{e.g.}$, users commenting on live streams) on platforms like Twitch exert additional pressures on the latency expected of such moderation systems. Despite their prevalence, relatively little is known about the effectiveness of these systems. In this paper, we conduct an audit of Twitch's automated moderation tool ($\texttt{AutoMod}$) to investigate its effectiveness in flagging hateful content. For our audit, we create streaming accounts to act as siloed test beds, and interface with the live chat using Twitch's APIs to send over $107,000$ comments collated from $4$ datasets. We measure $\texttt{AutoMod}$'s accuracy in flagging blatantly hateful content containing misogyny, racism, ableism and homophobia. Our experiments reveal that a large fraction of hateful messages, up to $94\%$ on some datasets, $\textit{bypass moderation}$. Contextual addition of slurs to these messages results in $100\%$ removal, revealing $\texttt{AutoMod}$'s reliance on slurs as a moderation signal. We also find that contrary to Twitch's community guidelines, $\texttt{AutoMod}$ blocks up to $89.5\%$ of benign examples that use sensitive words in pedagogical or empowering contexts. Overall, our audit points to large gaps in $\texttt{AutoMod}$'s capabilities and underscores the importance for such systems to understand context effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。