用大模型辅助撰写反仇恨与假信息并存内容的反驳文案
Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation

- 结合事实核查与非营利组织指南混合生成反驳文本
- 专家修改后自然度、完整性和指南遵循度显著提升
- 适合内容安全、公共传播领域研究者参考
仇恨言论与虚假信息常在网络中同时出现,加剧偏见与社会分裂。面对规模庞大,利用大语言模型(LLM)辅助专业反驳文案(CS)写作日益受到关注,但现有研究多将其分开处理。本文首次在两者共现情境下研究反驳文案生成,测试三种知识驱动策略:一是使用事实核查指南与文章;二是采用非营利组织(NGO)指南与报告;三是融合两方资源的混合策略。23位专家对生成内容进行修订,并通过人工与自动指标评估。结果显示,尽管模型在40%情况下生成可接受的反驳文案,但专家修改显著提升了自然度、完整性与指南符合度。基于修订后结果,混合策略在众包评估中表现最优,兼具强事实纠正、刻板印象缓解与共情互动。论文发布包含仇恨与虚假主张及专家验证反驳文案和支撑知识的数据集。
原文摘要 · Abstract (English)
Hate speech and misinformation frequently co-occur online, amplifying prejudice and polarization. Given their scale, using Large Language Models (LLMs) to assist expert counterspeech (CS) writing has gained interest, yet prior work has addressed these phenomena separately. We bridge this gap by studying CS generation in contexts where both hate and misinformation co-occur. We test three knowledge-driven generation strategies: first we prompt an LLM with fact-checkers' guidelines and fact-checking articles; secondly, with NGOs' guidelines and reports; thirdly, we create a mixed strategy that combines guidelines and documents from both. 23 experts revise the generated CS, which are assessed via human and automatic metrics. While LLMs produce adequate CS in 40% of cases, expert edits substantially improve naturalness, exhaustiveness, and adherence to guidelines. Based on the post-edited CS, the mixed strategy proves to be the most effective in crowdsourcing evaluation, pairing strong factual correction with stereotype mitigation and empathetic engagement. We release a dataset of hateful and misinformed claims with expert-verified CS and supporting knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。