构建多智能体大模型安全测试基准,揭示系统脆弱性。
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
- 设计五类场景300个攻击实例,涵盖六种攻击类型与211个工具。
- 多智能体系统在攻击下表现极差,多数模型任务成功率低于20%。
- 适合研究安全、防御或部署多智能体系统的团队参考。
大型语言模型(LLMs)通过工具使用、规划和决策能力展现出作为自主代理的强大潜力,被广泛应用于复杂任务。随着任务复杂度提升,多智能体LLM系统日益用于协同解决问题。然而,这类系统的安全与可靠性仍缺乏深入研究。现有评估基准多集中于单智能体场景,未能捕捉多智能体动态与协作带来的独特风险。为此,我们提出TAMAS(Threats and Attacks in Multi-Agent Systems),一个用于评估多智能体LLM系统鲁棒性与安全性的基准。TAMAS包含五个不同场景,共300个对抗实例,覆盖六种攻击类型和211个工具,以及100个无害任务。我们在十种主流LLM及Autogen和CrewAI框架中的三种交互配置下评估系统性能,揭示了当前多智能体部署中的关键挑战与失效模式。此外,我们引入有效鲁棒性评分(ERS),衡量安全性与任务有效性之间的权衡。结果表明,多智能体系统极易受对抗攻击影响,亟需更强的防御机制。TAMAS为系统化研究与提升多智能体LLM安全性提供了基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated strong capabilities as autonomous agents through tool use, planning, and decision-making abilities, leading to their widespread adoption across diverse tasks. As task complexity grows, multi-agent LLM systems are increasingly used to solve problems collaboratively. However, safety and security of these systems remains largely under-explored. Existing benchmarks and datasets predominantly focus on single-agent settings, failing to capture the unique vulnerabilities of multi-agent dynamics and co-ordination. To address this gap, we introduce $\textbf{T}$hreats and $\textbf{A}$ttacks in $\textbf{M}$ulti-$\textbf{A}$gent $\textbf{S}$ystems ($\textbf{TAMAS}$), a benchmark designed to evaluate the robustness and safety of multi-agent LLM systems. TAMAS includes five distinct scenarios comprising 300 adversarial instances across six attack types and 211 tools, along with 100 harmless tasks. We assess system performance across ten backbone LLMs and three agent interaction configurations from Autogen and CrewAI frameworks, highlighting critical challenges and failure modes in current multi-agent deployments. Furthermore, we introduce Effective Robustness Score (ERS) to assess the tradeoff between safety and task effectiveness of these frameworks. Our findings show that multi-agent systems are highly vulnerable to adversarial attacks, underscoring the urgent need for stronger defenses. TAMAS provides a foundation for systematically studying and improving the safety of multi-agent LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。