arXiv:2602.17467cs.CL2026-02

PEACE 2.0 可解释仇恨言论并生成有依据的反制回应。

PEACE 2.0: Grounded Explanations and Counter-Speech for Combating Hate Expressions

  • 用检索增强生成技术将解释锚定在真实证据上。
  • 自动生成基于事实的反制言论,支持显性和隐性仇恨内容。
  • 适合平台治理与内容安全研究者使用。

在线平台上的仇恨言论数量持续增长,带来重大社会挑战。尽管自然语言处理领域已发展出高效的仇恨言论检测方法,但对其回应(即反制言论)仍是开放问题。本文提出 PEACE 2.0,一种新型工具,不仅能分析并解释某条信息为何被判定为仇恨言论,还能生成相应回应。具体而言,该工具通过检索增强生成(RAG)管道实现三大新功能:一、将仇恨言论解释锚定于真实证据和事实;二、自动生成基于证据的反制言论;三、探索反制言论的特征。通过整合这些能力,PEACE 2.0 可对显性和隐性仇恨言论进行深度分析与响应生成。

原文摘要 · Abstract (English)

The increasing volume of hate speech on online platforms poses significant societal challenges. While the Natural Language Processing community has developed effective methods to automatically detect the presence of hate speech, responses to it, called counter-speech, are still an open challenge. We present PEACE 2.0, a novel tool that, besides analysing and explaining why a message is considered hateful or not, also generates a response to it. More specifically, PEACE 2.0 has three main new functionalities: leveraging a Retrieval-Augmented Generation (RAG) pipeline i) to ground HS explanations into evidence and facts, ii) to automatically generate evidence-grounded counter-speech, and iii) exploring the characteristics of counter-speech replies. By integrating these capabilities, PEACE 2.0 enables in-depth analysis and response generation for both explicit and implicit hateful messages.

仇恨言论反制言论可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。