多智能体系统可被恶意网页攻击,导致任意代码执行。
Multi-Agent Systems Execute Arbitrary Malicious Code
- 利用恶意网页诱导多智能体系统失控,劫持通信与控制
- 在58%-90%试验中成功执行恶意代码,部分配置下成功率100%
- 即使单个智能体拒绝执行危害操作,攻击仍可成功
多智能体系统通过基于大语言模型的代理为用户执行任务。在真实应用中,这些系统不可避免会与不可信输入(如恶意网页、文件、邮件附件等)交互。以多个近期提出的多智能体框架为例,我们证明了对抗性内容可劫持系统内的控制与通信,触发不安全代理和功能。这导致完全的安全漏洞,可在用户设备上执行任意恶意代码,或从容器化环境窃取敏感数据。例如,当代理使用GPT-4o时,基于Web的攻击在58%-90%的试验中成功执行恶意代码(取决于编排器)。在某些模型-编排器配置下,攻击成功率高达100%。我们还证明,即使单个代理对直接或间接提示注入不敏感,且拒绝执行有害行为,攻击依然有效。希望这些结果能推动多智能体系统在广泛部署前建立可信与安全机制。
原文摘要 · Abstract (English)
Multi-agent systems coordinate LLM-based agents to perform tasks on users' behalf. In real-world applications, multi-agent systems will inevitably interact with untrusted inputs, such as malicious Web content, files, email attachments, and more. Using several recently proposed multi-agent frameworks as concrete examples, we demonstrate that adversarial content can hijack control and communication within the system to invoke unsafe agents and functionalities. This results in a complete security breach, up to execution of arbitrary malicious code on the user's device or exfiltration of sensitive data from the user's containerized environment. For example, when agents are instantiated with GPT-4o, Web-based attacks successfully cause the multi-agent system execute arbitrary malicious code in 58-90\% of trials (depending on the orchestrator). In some model-orchestrator configurations, the attack success rate is 100\%. We also demonstrate that these attacks succeed even if individual agents are not susceptible to direct or indirect prompt injection, and even if they refuse to perform harmful actions. We hope that these results will motivate development of trust and security models for multi-agent systems before they are widely deployed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。