arXiv:2509.01379cs.CL2025-09被引 3

AI聊天助手辅助识别仇恨言论并解释原因,提升审核透明度。

WATCHED: A Web AI Agent Tool for Combating Hate Speech by Expanding Data

  • 用大模型结合多种工具分析内容与政策依据
  • 在测试中达到0.91的宏平均F1分数
  • 适合审核员、安全团队和研究者使用

在线危害是数字空间中的重大问题,威胁用户安全并削弱对社交媒体的信任。仇恨言论是最顽固的形式之一。为应对这一挑战,需要兼具自动化系统速度与人类审核员判断力的工具。这些工具不仅要识别有害内容,还需清晰解释决策依据,以建立信任。本文提出WATCHED,一个专为内容审核设计的聊天机器人,作为人工智能代理系统,利用大语言模型及多个专用工具:将新帖子与真实仇恨言论和中性内容对比,使用基于BERT的分类器标记有害信息,借助Urban Dictionary等资源解析俚语,生成思维链推理,并核查平台指南以支持决策。该组合使系统不仅能检测仇恨言论,还能基于先例与政策说明其判定理由。实验结果表明,该方法超越现有最佳水平,宏观F1得分为0.91。本工具面向审核员、安全团队及研究人员,通过人机协作助力降低网络危害。

原文摘要 · Abstract (English)

Online harms are a growing problem in digital spaces, putting user safety at risk and reducing trust in social media platforms. One of the most persistent forms of harm is hate speech. To address this, we need tools that combine the speed and scale of automated systems with the judgment and insight of human moderators. These tools should not only find harmful content but also explain their decisions clearly, helping to build trust and understanding. In this paper, we present WATCHED, a chatbot designed to support content moderators in tackling hate speech. The chatbot is built as an Artificial Intelligence Agent system that uses Large Language Models along with several specialised tools. It compares new posts with real examples of hate speech and neutral content, uses a BERT-based classifier to help flag harmful messages, looks up slang and informal language using sources like Urban Dictionary, generates chain-of-thought reasoning, and checks platform guidelines to explain and support its decisions. This combination allows the chatbot not only to detect hate speech but to explain why content is considered harmful, grounded in both precedent and policy. Experimental results show that our proposed method surpasses existing state-of-the-art methods, reaching a macro F1 score of 0.91. Designed for moderators, safety teams, and researchers, the tool helps reduce online harms by supporting collaboration between AI and human oversight.

AI审核仇恨言论大模型应用人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。