arXiv:2604.07775cs.AIcs.CL2026-04ACL被引 6

构建统一评估框架,检测多智能体系统中的恶意指令传播漏洞。

ACIArena: Toward Unified Evaluation for Agent Cascading Injection

  • 设计统一框架,覆盖外部输入、代理信息、消息交互等攻击面。
  • 提供1356个测试用例,验证不同角色与交互模式对系统鲁棒性的影响。
  • 揭示简单防御在真实场景中失效,强调角色设计与交互控制的重要性。

协作与信息共享虽增强多智能体系统(MAS)能力,却也引入关键安全风险——代理级联注入(ACI)。此类攻击中,被攻破的代理利用代理间信任传播恶意指令,引发系统级联失败。现有研究仅覆盖有限攻击策略与简化场景,泛化性不足。为此,我们提出ACIArena,一个统一的MAS鲁棒性评估框架。该框架涵盖外部输入、代理配置、跨代理消息等多维度攻击面,以及指令劫持、任务破坏、信息窃取等攻击目标。其统一规范支持MAS构建与攻防模块集成,覆盖六种主流实现,并提供1356个测试用例以系统评估鲁棒性。实验表明,仅依赖拓扑结构评估不足以衡量鲁棒性;鲁棒的MAS需精心设计角色与受控交互模式。此外,简化环境下的防御常无法迁移至真实场景,局部防御甚至引入新漏洞。ACIArena旨在为深入探索MAS设计原则提供坚实基础。

原文摘要 · Abstract (English)

Collaboration and information sharing empower Multi-Agent Systems (MAS) but also introduce a critical security risk known as Agent Cascading Injection (ACI). In such attacks, a compromised agent exploits inter-agent trust to propagate malicious instructions, causing cascading failures across the system. However, existing studies consider only limited attack strategies and simplified MAS settings, limiting their generalizability and comprehensive evaluation. To bridge this gap, we introduce ACIArena, a unified framework for evaluating the robustness of MAS. ACIArena offers systematic evaluation suites spanning multiple attack surfaces (i.e., external inputs, agent profiles, inter-agent messages) and attack objectives (i.e., instruction hijacking, task disruption, information exfiltration). Specifically, ACIArena establishes a unified specification that jointly supports MAS construction and attack-defense modules. It covers six widely used MAS implementations and provides a benchmark of 1,356 test cases for systematically evaluating MAS robustness. Our benchmarking results show that evaluating MAS robustness solely through topology is insufficient; robust MAS require deliberate role design and controlled interaction patterns. Moreover, defenses developed in simplified environments often fail to transfer to real-world settings; narrowly scoped defenses may even introduce new vulnerabilities. ACIArena aims to provide a solid foundation for advancing deeper exploration of MAS design principles.

多智能体安全评估攻防对抗鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。