arXiv:2510.12316cs.CL2025-10被引 2

用事实增强的对话生成系统,有效对抗针对八大群体的仇恨言论。

Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation

  • 基于检索增强生成,从权威文献库中调取事实支持反驳内容
  • 在8类目标群体上生成内容,人类评估与自动指标均优于主流模型
  • 适合反仇恨内容、事实核查及可信AI生成场景

对抗性言论生成是事实核查和仇恨言论应对中的核心任务。现有方法多依赖大语言模型或非政府组织专家,存在生成内容可靠性差、可扩展性不足等问题。为此,本文提出一种基于知识的对抗性言论生成框架,整合先进检索增强生成(RAG)流程,针对仇恨言论研究中识别出的8大目标群体(包括女性、有色人种、残障人士、移民、穆斯林、犹太人、LGBT群体等)生成可信反驳内容。构建了涵盖联合国数字图书馆、EUR-Lex及欧盟基本权利机构文件的32,792篇文本的知识库。通过MultiTarget-CONAN数据集进行评估,结合标准指标(JudgeLM)与人工评测,结果表明该框架在两项评估中均优于主流大模型基线与竞争方法。该框架与知识库为仇恨言论及其他场景下的可信、高质量对抗性言论生成提供了新路径。

原文摘要 · Abstract (English)

Counter-speech generation is at the core of many expert activities, such as fact-checking and hate speech, to counter harmful content. Yet, existing work treats counter-speech generation as pure text generation task, mainly based on Large Language Models or NGO experts. These approaches show severe drawbacks due to the limited reliability and coherence in the generated countering text, and in scalability, respectively. To close this gap, we introduce a novel framework to model counter-speech generation as knowledge-wise text generation process. Our framework integrates advanced Retrieval-Augmented Generation (RAG) pipelines to ensure the generation of trustworthy counter-speech for 8 main target groups identified in the hate speech literature, including women, people of colour, persons with disabilities, migrants, Muslims, Jews, LGBT persons, and other. We built a knowledge base over the United Nations Digital Library, EUR-Lex and the EU Agency for Fundamental Rights, comprising a total of 32,792 texts. We use the MultiTarget-CONAN dataset to empirically assess the quality of the generated counter-speech, both through standard metrics (i.e., JudgeLM) and a human evaluation. Results show that our framework outperforms standard LLM baselines and competitive approach, on both assessments. The resulting framework and the knowledge base pave the way for studying trustworthy and sound counter-speech generation, in hate speech and beyond.

对抗性言论RAG事实核查仇恨言论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。