首个多轮跨语言仇恨与谣言对抗对话数据集,支持事实溯源。
CATCH-ME if you RAG: a dataset of Contextually Annotated multi-Turn Counterspeech against Hate and Misinformation Exchanges

- 构建多轮、多语言、带事实标注的对抗性对话数据集
- 覆盖5种语言,涵盖7类边缘群体,含文档级事实锚点
- 专为RAG系统设计,助力生成有据可查的反谣言回应
在线仇恨言论与虚假信息常相互交织,但现有NLP研究多将其孤立处理。尽管大模型可规模化辅助生成反制言论,零样本模型常产生重复模糊的回应,亟需高质量示例引导生成。然而,现有针对仇恨与谣言交叉场景的反制对话数据集稀缺,且仅限单轮英文对话,无法反映真实多轮多语交流。为此,我们首次推出大规模、专家标注、多语言的对话数据集,聚焦仇恨与谣言的交集问题。为确保事实准确性,对话均基于经验证的外部知识(如事实核查文章与非政府组织报告),并包含文档级和片段级的指代标注,可直接用于检索增强生成(RAG)系统。该数据集覆盖五种语言,针对七类边缘群体的仇恨言论,为训练与评估更具说服力、事实可靠的反制模型提供新资源。
原文摘要 · Abstract (English)
Online hate speech and misinformation frequently overlap, yet NLP research has mainly treated them in isolation. While LLMs represent a scalable solution for assisting humans in the generation of counterspeech for both threats, zero-shot models frequently generate repetitive and vague responses, underscoring the need for high-quality examples to steer model generation. However, existing counterspeech datasets against the overlap of hate and misinformation are scarce and limited to single-turn English dialogues, while real-life interactions span across multiple turns and languages. To bridge this gap, we introduce the first large-scale, expert-curated, multilingual dataset of dialogues tackling the intersection of hate and misinformation. To ensure factual grounding, the dialogues are also anchored in verified external knowledge (i.e., fact-checking articles and NGO reports) and include document- and chunk-level span annotations, making it directly applicable for RAG systems. Covering five languages and targeting hate directed at seven marginalized groups, this novel resource enables the training and evaluation of more persuasive, factually grounded counterspeech models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。