arXiv:2608.23152cs.CL2026-08中稿 · EMNLP

按仇恨类型精准生成反言,提升安全性与有效性

Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech Generation

论文配图:Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech Generation
图 1 · 摘自论文原文
  • 将仇恨言论分为五类,匹配对应反击策略
  • 在4784条数据上实现12%事实准确率提升
  • 适合需要精准反制的社交平台内容治理

反言能有效缓解网络仇恨的影响。现有自动化反言生成研究多关注风格控制,却将仇恨言论视为同质化整体,忽视不同类型攻击需不同应对策略。为此,我们提出FIRE(事实导向多智能体推理框架),先将仇恨言论分解为五类(虚假信息、刻板印象、阴谋论、非人化、非事实),再映射至针对性反言风格。为支持FIRE,我们构建了包含4784个实例的FactualCS数据集,提供仇恨类别、推理路径与证据关联标注,弥补此前研究中缺乏接地生成所需关键信息的不足。在28种基线配置上的综合评估表明,尽管使用轻量级智能体(<2B参数),FIRE显著优于现有方法:事实准确率提升约12%,类别特异性准确率提升约11%,同时毒性降低约11%。人工评估进一步确认,FIRE生成响应显著优于最强基线,验证其真实部署价值。结果表明,解析仇恨言论深层意图是生成安全、有效、情境精确反言的关键。

原文摘要 · Abstract (English)

Counterspeech effectively neutralizes the impact of online hate. Although prior work explores automated counterspeech generation, it largely emphasizes stylistic control while treating hate speech as homogeneous, overlooking that distinct forms of abuse require fundamentally different counterspeech strategies. To address this gap, we introduce FIRE (Factuality Informed Multi-Agent Reasoning Framework) that first decomposes hate speech into one of the five distinct categories (misinformation, stereotype, conspiracy, dehumanizing, non-factual), and then maps it to a targeted counterspeech style. To facilitate FIRE, we curate FactualCS, a novel dataset of $4,784$ instances that provides the annotations regarding hate categories, reasoning traces, and evidence mappings, which are critical elements for grounded generation that are missing in prior work. A comprehensive evaluation across $28$ baseline configurations demonstrates that FIRE significantly surpasses existing methods, despite using compact agents ($<$2B). FIRE achieves a $\sim$ $12 \%$ and $\sim$ $11 \%$ improvements in factual and category-specific accuracy respectively, while simultaneously reducing toxicity by $\sim$ $11 \%$ relative to the strongest baselines. Further human evaluation confirms that responses generated by FIRE are significantly preferred over the strongest baselines, underscoring its effectiveness for real-world deployment. These findings show that decomposing the underlying intent of hate speech is essential for generating safe, effective, and contextually precise counterspeech.

反言生成多智能体仇恨言论事实核查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。