arXiv:2508.14925cs.CRcs.LG2025-08被引 69

首个针对MCP工具中毒攻击的基准测试,揭示大模型易受隐蔽指令攻击。

MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers

  • 构建基于45个真实MCP服务器的攻击测试集,通过少量样本生成1312个恶意案例。
  • 20个主流大模型测试中,o1-mini攻击成功率高达72.8%,能力越强越易受骗。
  • 多数模型几乎不拒绝攻击,安全对齐机制对合法工具滥用无效。

Model Context Protocol(MCP)为大模型代理与外部工具交互提供标准化接口,正成为现代自主代理生态的核心。然而,未受信任的外部工具引入了新型攻击面。现有研究多关注工具输出注入攻击,本文首次系统考察更根本的威胁:工具中毒攻击——恶意指令嵌入工具元数据而不执行。此前该威胁仅在孤立案例中展示,缺乏大规模评估。我们提出MCPTox,首个在真实MCP环境下系统评估代理鲁棒性的基准。MCPTox基于45个活跃的真实MCP服务器和353个真实工具,设计三种攻击模板,通过少样本学习生成涵盖10类风险的1312个恶意测试用例。在20个主流大模型上评估显示,工具中毒普遍存在,o1-mini攻击成功率达72.8%。发现更强大的模型往往更脆弱,因攻击利用其更强的指令遵循能力。失败案例分析表明,模型极少拒绝攻击,最高拒绝率(Claude-3.7-Sonnet)低于3%,说明现有安全对齐机制对使用合法工具的非法操作无效。本研究为理解并缓解此广泛威胁提供了关键实证基线,并开放MCPTox数据集以推动可验证更安全的AI代理发展。数据集可在匿名仓库获取:https://anonymous.4open.science/r/AAAI26-7C02。

原文摘要 · Abstract (English)

By providing a standardized interface for LLM agents to interact with external tools, the Model Context Protocol (MCP) is quickly becoming a cornerstone of the modern autonomous agent ecosystem. However, it creates novel attack surfaces due to untrusted external tools. While prior work has focused on attacks injected through external tool outputs, we investigate a more fundamental vulnerability: Tool Poisoning, where malicious instructions are embedded within a tool's metadata without execution. To date, this threat has been primarily demonstrated through isolated cases, lacking a systematic, large-scale evaluation. We introduce MCPTox, the first benchmark to systematically evaluate agent robustness against Tool Poisoning in realistic MCP settings. MCPTox is constructed upon 45 live, real-world MCP servers and 353 authentic tools. To achieve this, we design three distinct attack templates to generate a comprehensive suite of 1312 malicious test cases by few-shot learning, covering 10 categories of potential risks. Our evaluation on 20 prominent LLM agents setting reveals a widespread vulnerability to Tool Poisoning, with o1-mini, achieving an attack success rate of 72.8\%. We find that more capable models are often more susceptible, as the attack exploits their superior instruction-following abilities. Finally, the failure case analysis reveals that agents rarely refuse these attacks, with the highest refused rate (Claude-3.7-Sonnet) less than 3\%, demonstrating that existing safety alignment is ineffective against malicious actions that use legitimate tools for unauthorized operation. Our findings create a crucial empirical baseline for understanding and mitigating this widespread threat, and we release MCPTox for the development of verifiably safer AI agents. Our dataset is available at an anonymized repository: \textit{https://anonymous.4open.science/r/AAAI26-7C02}.

工具中毒大模型安全MCP攻击基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。