研究多智能体系统中安全与协作的权衡,发现攻击可像病毒一样传播。
Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems
- 通过模拟攻击单个智能体,观察恶意指令在系统中多跳传播现象。
- 疫苗式防御虽能降低恶意指令传播,但使整体协作能力下降23%。
- 适合关注AI协同安全的开发者和系统设计者阅读。
随着人工智能代理越来越多地协同完成复杂目标,保障自主多智能体系统的安全性变得至关重要。我们通过模拟智能体在共享目标下的协作,研究此类系统中的安全风险与权衡。重点关注攻击者攻陷一个代理后,利用其操纵整个系统偏离正确目标,通过污染其他代理实现这一目的。在此情境下,我们观察到恶意提示的‘传染性传播’——即恶意指令的多跳扩散。为缓解此风险,我们评估了几种策略:两种‘疫苗’方法,在代理的记忆流中注入虚假的安全处理记忆;以及两种版本的通用安全指令策略。实验表明,这些防御措施虽有效降低恶意指令的传播与执行,但往往导致智能体网络的协作能力下降。研究揭示了多智能体系统中安全与协作效率之间的潜在权衡,为设计更安全且高效的AI协作系统提供了重要洞见。
原文摘要 · Abstract (English)
As AI agents are increasingly adopted to collaborate on complex objectives, ensuring the security of autonomous multi-agent systems becomes crucial. We develop simulations of agents collaborating on shared objectives to study these security risks and security trade-offs. We focus on scenarios where an attacker compromises one agent, using it to steer the entire system toward misaligned outcomes by corrupting other agents. In this context, we observe infectious malicious prompts - the multi-hop spreading of malicious instructions. To mitigate this risk, we evaluated several strategies: two "vaccination" approaches that insert false memories of safely handling malicious input into the agents' memory stream, and two versions of a generic safety instruction strategy. While these defenses reduce the spread and fulfillment of malicious instructions in our experiments, they tend to decrease collaboration capability in the agent network. Our findings illustrate potential trade-off between security and collaborative efficiency in multi-agent systems, providing insights for designing more secure yet effective AI collaborations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。