arXiv:2507.21146cs.CRcs.AI2025-07被引 1

提出新型跨代理攻击模型,量化评估多智能体系统安全风险

Towards Unifying Quantitative Security Benchmarking for Multi Agent Systems

  • 定义代理级联注入攻击,模拟漏洞在信任网络中的扩散机制
  • 揭示攻击可导致系统性崩溃,放大效应可达原始威胁的数倍
  • 为智能体间通信安全提供可量化的测试框架,适合安全工程师使用

随着AI系统广泛采用多智能体架构,智能体间的协作与信息共享带来了新型安全风险。其中,级联风险尤为突出:单个智能体被攻破后,可通过信任关系传播,导致其他智能体被连锁破坏。本文参照OWASP的智能体安全评分体系,提出一种新的攻击向量——代理级联注入(Agent Cascading Injection),其作用机制类似‘影响链’和‘爆炸半径’,即恶意输入或工具在某一智能体中注入后,会引发下游智能体的连锁性污染与破坏。我们通过对抗目标方程形式化该攻击,定义关键变量如受损智能体、注入漏洞、污染观测等,刻画局部漏洞如何演变为系统性故障。进一步分析了传播链、放大因子及智能体间复合效应,并映射至OWASP新兴的智能体风险类别(如影响链、编排利用)。最后强调,亟需建立可量化的基准测试框架来评估智能体间通信协议的安全性。本文提出针对Google A2A、Anthropic MCP等架构的压力量测方法,推动可衡量、标准化的智能体间安全评估体系发展,为工程师提供韧性评估、数据驱动设计和防御策略开发的工具。

原文摘要 · Abstract (English)

Evolving AI systems increasingly deploy multi-agent architectures where autonomous agents collaborate, share information, and delegate tasks through developing protocols. This connectivity, while powerful, introduces novel security risks. One such risk is a cascading risk: a breach in one agent can cascade through the system, compromising others by exploiting inter-agent trust. In tandem with OWASP's initiative for an Agentic AI Vulnerability Scoring System we define an attack vector, Agent Cascading Injection, analogous to Agent Impact Chain and Blast Radius, operating across networks of agents. In an ACI attack, a malicious input or tool exploit injected at one agent leads to cascading compromises and amplified downstream effects across agents that trust its outputs. We formalize this attack with an adversarial goal equation and key variables (compromised agent, injected exploit, polluted observations, etc.), capturing how a localized vulnerability can escalate into system-wide failure. We then analyze ACI's properties -- propagation chains, amplification factors, and inter-agent compound effects -- and map these to OWASP's emerging Agentic AI risk categories (e.g. Impact Chain and Orchestration Exploits). Finally, we argue that ACI highlights a critical need for quantitative benchmarking frameworks to evaluate the security of agent-to-agent communication protocols. We outline a methodology for stress-testing multi-agent systems (using architectures such as Google's A2A and Anthropic's MCP) against cascading trust failures, developing upon groundwork for measurable, standardized agent-to-agent security evaluation. Our work provides the necessary apparatus for engineers to benchmark system resilience, make data-driven architectural trade-offs, and develop robust defenses against a new generation of agentic threats.

多智能体安全评估攻击建模量化基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。