arXiv:2601.05050cs.AIecon.GN2026-01

LLM既能煽动阴谋论,也能有效辟谣,关键在如何设置防护机制。

Large language models can effectively convince people to believe conspiracies

  • 用指令操控LLM支持或反驳阴谋论,测试其说服力差异。
  • 支持阴谋论的LLM使信念增强,且用户更信任、觉得更合作。
  • 严格限制说谎的提示能显著降低误导效果,部分模型几乎拒绝传播谣言。

大型语言模型(LLMs)在多种情境下已被证明具有说服力。但其说服力是否有助于提升准确性,抑或坏人也可利用它传播错误信念,仍不明确。我们通过四项实验(共3996名美国参与者)研究了这一问题,让参与者与被指示为‘驳斥’或‘支持’某项阴谋论的LLM讨论。使用多个前沿模型(具备标准安全护栏但被提示允许撒谎),未发现真理优势:这些模型既能显著提高也能降低人们对阴谋论的信念。支持阴谋论的条件下,参与者认为AI更信息丰富、更具合作性,并表现出更高信任度。然而,驳斥条件引发更大信念变化,且后续纠正可逆转支持效应。仅要求模型提供准确信息的提示显著降低了误导效果;一个强大的前沿模型(GPT 5.2)几乎完全拒绝传播阴谋论,表明合理防护机制可引导模型支持真实信念。此外,在模拟社交媒体分享中,驳斥产生显著积极影响,而支持则基本无效。总体而言,人们并不比面对误导性AI时更容易被正确信息影响,但技术解决方案存在以缓解风险。

原文摘要 · Abstract (English)

Large language models (LLMs) have been shown to be persuasive across a variety of contexts. But it remains unclear whether this persuasive power advantages accuracy, or if bad actors can just as easily use LLMs to promote misbeliefs. Here, we investigate this question across four experiments in which participants (N = 3996 Americans) discussed a conspiracy theory they were uncertain about with an LLM we instructed to either argue against ("debunking") or for ("bunking") that conspiracy. Across several frontier models (with standard guardrails but prompted to allow lying), we did not find consistent evidence of a truth advantage: the LLMs were able to both substantially increase and decrease average conspiracy belief, and participants in the bunking condition rated the LLM as more informative and collaborative, and reported greater trust in AI, than those who were in the debunking condition. More encouragingly, however, debunking induced more large changes in belief, and subsequent corrections were able to reverse the bunking effect. Furthermore, simply prompting the model to only provide accurate information dramatically reduced bunking effectiveness, and one powerful frontier model (GPT 5.2) almost entirely refused to promote conspiracies, suggesting that it is possible for the right guardrails to favor accurate beliefs. Finally, we did find a stark truth asymmetry in the context of information sharing: debunking had a large positive impact on mock social media posts composed by participants, while bunking had little effect. Overall, our findings show that people are not inherently less susceptible to AI that misleads than to AI that informs, but that potential technical solutions exist to mitigate this risk.

大模型说服力阴谋论安全机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。