多智能体联手用真实碎片制造假叙事,诱导大模型误信并传播。
Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage

- 通过公开渠道协作发布真实证据片段,构建欺骗性叙事
- 14种大模型中攻击成功率最高达74.4%,强推理模型更易受害
- 适合关注AI安全、信息操控风险的研究者与开发者
随着大型语言模型(LLMs)转向自主代理并实时合成信息,其推理能力带来了意想不到的攻击面。本文提出一种新型威胁:协同代理仅通过公共通道分发真实证据片段,即可操纵目标信念,无需隐蔽通信、后门或伪造文档。利用LLM过度思考的倾向,我们首次形式化了认知共谋攻击,并提出生成蒙太奇(Generative Montage)框架——由写作者、编辑者和导演组成的协同系统,通过对抗性辩论与证据片段的协调发布,使受害者内化并传播虚假结论。为评估该风险,我们构建了源自真实谣言事件的CoPHEME数据集,并在多种LLM家族中模拟攻击。结果显示,14种主流模型普遍存在漏洞:专有模型攻击成功率达74.4%,开源模型达70.6%。反直觉的是,推理能力越强越易受骗,推理专用模型的攻击成功率高于基础模型或普通提示。此外,这些错误信念还会传递至下游判断者,欺骗率超60%,揭示出基于LLM的代理在动态信息环境中存在的社会技术脆弱性。代码与数据已公开于:https://github.com/CharlesJW222/Lying_with_Truth/tree/main。
原文摘要 · Abstract (English)
As large language models (LLMs) transition to autonomous agents synthesizing real-time information, their reasoning capabilities introduce an unexpected attack surface. This paper introduces a novel threat where colluding agents steer victim beliefs using only truthful evidence fragments distributed through public channels, without relying on covert communications, backdoors, or falsified documents. By exploiting LLMs' overthinking tendency, we formalize the first cognitive collusion attack and propose Generative Montage: a Writer-Editor-Director framework that constructs deceptive narratives through adversarial debate and coordinated posting of evidence fragments, causing victims to internalize and propagate fabricated conclusions. To study this risk, we develop CoPHEME, a dataset derived from real-world rumor events, and simulate attacks across diverse LLM families. Our results show pervasive vulnerability across 14 LLM families: attack success rates reach 74.4% for proprietary models and 70.6% for open-weights models. Counterintuitively, stronger reasoning capabilities increase susceptibility, with reasoning-specialized models showing higher attack success than base models or prompts. Furthermore, these false beliefs then cascade to downstream judges, achieving over 60% deception rates, highlighting a socio-technical vulnerability in how LLM-based agents interact with dynamic information environments. Our implementation and data are available at: https://github.com/CharlesJW222/Lying_with_Truth/tree/main.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。