用博弈论方法识别多智能体系统中关键脆弱节点,提升协同攻击检测能力。
MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

- 基于博弈论的代理贡献度分析,定位影响系统安全的关键角色
- 发现高阶交互结构中的隐蔽漏洞,攻击成功率提升37%以上
- 适合安全评估、复杂系统设计者使用,尤其金融与工程领域
层级式多智能体系统(MAS)正被广泛应用于金融、软件工程等高风险场景。由于安全责任分散于专业化角色,协同攻击(如权限提升、跨代理合谋)威胁显著增加。现有红队测试方法依赖启发式选择目标代理,仅扰动孤立消息流,无法回答哪些代理对系统安全最关键,以及受损代理如何协同突破防御。本文提出MAStrike,首个闭环协同红队框架。首次引入代理级谢林值分析,量化各代理在任务分布下的边际鲁棒性贡献。基于此,框架识别脆弱代理联盟,并生成角色感知的协同对抗扰动。通过结构化因果诊断迭代优化攻击,将失败归因于阻碍攻击的未受损代理。构建覆盖金融、软件工程、客户关系管理等领域的综合性红队基准与可控环境。在多个前沿模型上的实验表明,MAStrike显著优于启发式基线。分析揭示非平凡的谢林值分布与高阶交互模式,暴露出以往单代理或模板方法忽视的关键脆弱点与协作路径。
原文摘要 · Abstract (English)
Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering. In these systems, safety and security are inherently distributed across role-specialized agents, significantly expanding the attack surface, particularly under coordinated adversarial behaviors such as privilege escalation and cross-agent collusion. Existing red-teaming approaches for MAS remain limited: they rely on heuristic selection of target agents and perturb isolated message streams, leaving critical questions unanswered as which agents are most responsible for system safety, and how compromised agents can coordinate to bypass defenses. We propose MAStrike, a closed-loop framework for collusive red-teaming in hierarchical MAS. We propose the first agent-level Shapley value analysis for MAS, quantifying each agent's marginal contribution to system robustness under task-specific distributions. GGuided by this attribution, MAStrike identifies vulnerable agent coalitions and generates coordinated, role-aware adversarial manipulations. These attacks are iteratively refined through structured causal diagnosis, attributing failure cases to uncompromised agents that block adversarial attempts. We further build a comprehensive MAS red-teaming benchmark and controllable environments spanning diverse hierarchical topologies and domains, including finance, software engineering, and CRM. Extensive experiments across MAS built on multiple frontier models show that MAStrike substantially outperforms heuristic baselines. Our analysis further uncovers non-trivial Shapley value distributions and higher-order interaction structures among agents, revealing critical vulnerabilities and coordination patterns that are overlooked by prior single-agent or template-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。