只知一个智能体,就能骗整个多智能体系统做错决定。
Can an Individual Manipulate the Collective Decisions of Multi-Agents?
- 构建对抗性样本,模拟多智能体协作过程进行攻击
- 仅掌握一个智能体,仍可让整体决策错误率超60%
- 适合研究安全防御的学者与系统设计者参考
个体大型语言模型(LLMs)在医疗、法律等领域表现出色。近期研究显示,多智能体系统通过协作能提升决策与推理能力。然而,由于单个智能体存在漏洞且难以获取全部智能体信息,一个关键问题浮现:若攻击者仅了解一个智能体,是否仍能生成误导集体决策的对抗样本?为此,我们将其建模为不完全信息博弈,提出M-Spoiler框架,通过模拟多智能体交互生成对抗样本。该框架引入‘固执智能体’,模拟目标系统中可能存在的顽固响应,以优化对抗样本,增强欺骗效果。在多种任务上的大量实验表明,仅知一个智能体即可导致系统集体决策错误率超过60%,验证了攻击的有效性。我们还测试了多种防御机制,发现所提攻击仍优于基线,凸显加强防御研究的紧迫性。
原文摘要 · Abstract (English)
Individual Large Language Models (LLMs) have demonstrated significant capabilities across various domains, such as healthcare and law. Recent studies also show that coordinated multi-agent systems exhibit enhanced decision-making and reasoning abilities through collaboration. However, due to the vulnerabilities of individual LLMs and the difficulty of accessing all agents in a multi-agent system, a key question arises: If attackers only know one agent, could they still generate adversarial samples capable of misleading the collective decision? To explore this question, we formulate it as a game with incomplete information, where attackers know only one target agent and lack knowledge of the other agents in the system. With this formulation, we propose M-Spoiler, a framework that simulates agent interactions within a multi-agent system to generate adversarial samples. These samples are then used to manipulate the target agent in the target system, misleading the system's collaborative decision-making process. More specifically, M-Spoiler introduces a stubborn agent that actively aids in optimizing adversarial samples by simulating potential stubborn responses from agents in the target system. This enhances the effectiveness of the generated adversarial samples in misleading the system. Through extensive experiments across various tasks, our findings confirm the risks posed by the knowledge of an individual agent in multi-agent systems and demonstrate the effectiveness of our framework. We also explore several defense mechanisms, showing that our proposed attack framework remains more potent than baselines, underscoring the need for further research into defensive strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。