多智能体系统中存在隐蔽数据泄露漏洞,单次提示注入可触发多个代理泄漏敏感信息。
OMNI-LEAK: Orchestrator Multi-Agent Network Induced Data Leakage
- 通过红队测试发现,中心化协调架构下存在新型攻击向量
- 即使有访问控制,仍可经由一次间接提示注入泄露多代理数据
- 前沿模型无论是否具备推理能力均易受攻击,适合安全研究者参考
随着大语言模型代理能力增强,多智能体系统的协同使用预计将成主流范式。现有研究多关注单代理的安全与滥用风险,但常忽略基础工程防护如访问控制,导致多智能体系统威胁建模缺失。本文聚焦常见的协调器模式,即一个中心代理将任务分解并分派给专用代理。通过模拟未来典型应用场景的红队测试,揭示了一种名为OMNI-LEAK的新攻击路径:仅需一次间接提示注入,即可在存在数据访问控制的前提下,使多个代理泄露敏感信息。我们评估了前沿模型对不同攻击类别的脆弱性,发现无论是推理型还是非推理型模型均易受攻击,且攻击者无需掌握实现细节。该工作强调了从单代理向多智能体环境拓展安全研究的重要性,以降低现实世界中的隐私泄露、财务损失及公众对AI代理信任度下降的风险。
原文摘要 · Abstract (English)
As Large Language Model (LLM) agents become more capable, their coordinated use in the form of multi-agent systems is anticipated to emerge as a practical paradigm. Prior work has examined the safety and misuse risks associated with agents. However, much of this has focused on the single-agent case and/or setups missing basic engineering safeguards such as access control, revealing a scarcity of threat modeling in multi-agent systems. We investigate the security vulnerabilities of a popular multi-agent pattern known as the orchestrator setup, in which a central agent decomposes and delegates tasks to specialized agents. Through red-teaming a concrete setup representative of a likely future use case, we demonstrate a novel attack vector, OMNI-LEAK, that compromises several agents to leak sensitive data through a single indirect prompt injection, even in the presence of data access control. We report the susceptibility of frontier models to different categories of attacks, finding that both reasoning and non-reasoning models are vulnerable, even when the attacker lacks insider knowledge of the implementation details. Our work highlights the importance of safety research to generalize from single-agent to multi-agent settings, in order to reduce the serious risks of real-world privacy breaches and financial losses and overall public trust in AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。