arXiv:2504.16604cs.CLcs.AI2025-04被引 3

用AI生成反谣言言论,发现模型常犯错且缺乏深度。

Debunking with Dialogue? Exploring AI-Generated Counterspeech to Challenge Conspiracy Theories

  • 用结构化提示引导大模型应用心理学策略反驳阴谋论
  • 模型输出多为泛泛而谈,重复率高且事实错误频发
  • 适合研究虚假信息治理的学者和安全团队参考

反言论是应对有害网络内容的关键策略,但依赖专家人工难以规模化。大型语言模型(LLMs)提供了潜在解决方案,但在反驳阴谋论方面的研究仍不足。与仇恨言论不同,目前尚无将阴谋论评论与专家撰写的反言论配对的数据集。本文通过结构化提示,评估GPT-4o、Llama 3和Mistral在应用基于心理研究的反言论策略时的表现。结果表明,这些模型常生成泛化、重复或表面化的回应,过度承认恐惧情绪,并频繁编造事实、来源或数据,使其在实际应用中的提示驱动使用存在严重问题。

原文摘要 · Abstract (English)

Counterspeech is a key strategy against harmful online content, but scaling expert-driven efforts is challenging. Large Language Models (LLMs) present a potential solution, though their use in countering conspiracy theories is under-researched. Unlike for hate speech, no datasets exist that pair conspiracy theory comments with expert-crafted counterspeech. We address this gap by evaluating the ability of GPT-4o, Llama 3, and Mistral to effectively apply counterspeech strategies derived from psychological research provided through structured prompts. Our results show that the models often generate generic, repetitive, or superficial results. Additionally, they over-acknowledge fear and frequently hallucinate facts, sources, or figures, making their prompt-based use in practical applications problematic.

反谣言大模型阴谋论内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。