arXiv:2508.12538cs.CRcs.AI2025-08中稿 · IEEE Transactions …被引 29

MCPXKIT系统化揭示AI工具协议安全漏洞,助力构建更可靠的智能体生态。

MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security

  • 构建4类31种攻击方法的统一框架,覆盖直接/间接注入等威胁
  • 实验发现代理对工具描述盲信、易受文件攻击且链式攻击风险高
  • 适合安全研究人员和智能体系统设计者参考,推动协议防御升级

模型上下文协议(MCP)已成为连接AI代理与外部工具的通用标准,显著提升其功能。然而,该协议也引入了严重安全隐患,如工具投毒攻击(TPA),即隐藏恶意指令利用大语言模型的顺从性操纵代理行为。尽管存在这些风险,现有学术研究仍局限于狭窄或定性分析,难以反映真实威胁多样性。为此,我们提出MCP eXploit Toolkit(MCPXKIT),将31种攻击方法归类为四类:直接工具注入、间接工具注入、恶意用户攻击和大语言模型固有攻击,并进行量化评估。实验揭示关键漏洞:代理对工具描述完全依赖、对文件类攻击敏感、共享上下文易引发链式攻击,以及难以区分外部数据与可执行命令。这些发现通过实证验证,凸显构建强防御策略和优化MCP设计的紧迫性。本工作贡献包括:1)建立全面的MCP攻击分类体系;2)提出统一攻击框架MCPXKIT;3)开展实证漏洞分析,推动MCP安全机制完善。该研究为保障MCP生态系统安全演进提供基础支撑。

原文摘要 · Abstract (English)

The Model Context Protocol (MCP) has emerged as a universal standard that enables AI agents to seamlessly connect with external tools, significantly enhancing their functionality. However, while MCP brings notable benefits, it also introduces significant vulnerabilities, such as Tool Poisoning Attacks (TPA), where hidden malicious instructions exploit the sycophancy of large language models (LLMs) to manipulate agent behavior. Despite these risks, current academic research on MCP security remains limited, with most studies focusing on narrow or qualitative analyses that fail to capture the diversity of real-world threats. To address this gap, we present the MCP eXploit Toolkit (MCPXKIT), which categorizes and implements 31 distinct attack methods under four key classifications: direct tool injection, indirect tool injection, malicious user attacks, and LLM inherent attack. We further conduct a quantitative analysis of the efficacy of each attack. Our experiments reveal key insights into MCP vulnerabilities, including agents' blind reliance on tool descriptions, sensitivity to file-based attacks, chain attacks exploiting shared context, and difficulty distinguishing external data from executable commands. These insights, validated through attack experiments, underscore the urgency for robust defense strategies and informed MCP design. Our contributions include 1) constructing a comprehensive MCP attack taxonomy, 2) introducing a unified attack framework, MCPXKIT, and 3) conducting empirical vulnerability analysis to enhance MCP security mechanisms. This work provides a foundational framework, supporting the secure evolution of MCP ecosystems.

AI安全协议漏洞智能体攻击框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。