研究大模型多智能体系统中错误信息如何传播与缓解
Misinformation Propagation in Benign Multi-Agent Systems
- 通过注入意图误导信息,测试单/多智能体推理表现
- 错误信息在多智能体辩论中持续存在,但群体决策可减轻损失
- 群体构成与决策机制影响抗错能力,共识更稳定
多智能体系统通过大语言模型轮流交互解决复杂问题,正被应用于医疗诊断、法律分析等高风险场景。当单个智能体基于错误或误导性上下文(如工具调用结果)推理时,错误可能在交互中传播。本研究在推理、知识和对齐任务中向良性单智能体与多智能体系统注入基于意图的虚假信息。结果发现,错误信息会降低单智能体性能,并在多智能体辩论中持续存在,智能体常保留受误导同伴引入的答案。然而,相较于单智能体提示,多智能体辩论能减少性能下降,尤其当多数智能体未接触错误信息时。鲁棒性取决于群体构成与决策协议:在同侪压力下,共识比投票更稳定;多数意见通常可引导受误导智能体回归正确答案。结果表明,多智能体系统的抗错能力不仅依赖底层模型,还取决于信息交换方式与决策聚合机制。
原文摘要 · Abstract (English)
Multi-agent systems, in which multiple large language model agents solve problems through turn-based interaction, are increasingly deployed in high-stakes settings such as medical diagnosis, legal analysis, and forensic decision-making. Their reliability can be at risk when single agents reason from incorrect or misleading context, e.g., from tool calls, since errors may propagate through agent interactions. This work studies this risk by injecting intent-based misinformation into benign single-agent and multi-agent systems across reasoning, knowledge, and alignment tasks. We find that misinformation can degrade single-agent performance and persists across multi-agent debate, with agents often retaining answers introduced by misinformed peers. Nevertheless, multi-agent debate reduces the resulting performance degradation compared to single-agent prompting, especially when most agents are not exposed to misinformation. Robustness depends on group composition and decision protocol. Consensus can be more stable than voting under peer pressure, while majorities can often steer misinformed agents back toward correct answers. Our results show that misinformation robustness in multi-agent systems depends on the underlying model and also on how agents exchange information and aggregate decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。