评测大模型多智能体系统的抗攻击能力,发现规划器是弱点
PEAR: Planner-Executor Agent Robustness Benchmark
- 构建针对规划-执行结构的系统性评测基准
- 弱规划器比弱执行器更严重影响任务表现
- 攻击规划器最有效,性能与鲁棒性存在权衡
基于大语言模型的多智能体系统已成为解决跨领域复杂多步任务的强大范式。然而,尽管能力突出,多智能体系统仍易受对抗性操纵。现有研究多聚焦单一攻击面或特定场景,缺乏对系统漏洞的整体理解。为此,我们提出PEAR基准,用于系统评估规划-执行型多智能体系统的有效性与脆弱性。该基准兼容多种架构,但重点针对广泛采用的规划-执行结构。通过大量实验发现:(1) 弱规划器对整体干净任务性能的影响远大于弱执行器;(2) 规划器需要记忆模块,而执行器是否具备记忆模块对干净任务性能无显著影响;(3) 任务性能与鲁棒性之间存在权衡;(4) 针对规划器的攻击尤其有效,能有效误导系统。这些发现为提升多智能体系统鲁棒性提供了可操作建议,并为多智能体场景中的系统性防御奠定基础。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based Multi-Agent Systems (MAS) have emerged as a powerful paradigm for tackling complex, multi-step tasks across diverse domains. However, despite their impressive capabilities, MAS remain susceptible to adversarial manipulation. Existing studies typically examine isolated attack surfaces or specific scenarios, leaving a lack of holistic understanding of MAS vulnerabilities. To bridge this gap, we introduce PEAR, a benchmark for systematically evaluating both the utility and vulnerability of planner-executor MAS. While compatible with various MAS architectures, our benchmark focuses on the planner-executor structure, which is a practical and widely adopted design. Through extensive experiments, we find that (1) a weak planner degrades overall clean task performance more severely than a weak executor; (2) while a memory module is essential for the planner, having a memory module for the executor does not impact the clean task performance; (3) there exists a trade-off between task performance and robustness; and (4) attacks targeting the planner are particularly effective at misleading the system. These findings offer actionable insights for enhancing the robustness of MAS and lay the groundwork for principled defenses in multi-agent settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。