arXiv:2508.13220cs.CRcs.AI2025-08被引 53

首个针对模型上下文协议的系统化安全评测基准,揭示了通用接口的严重漏洞。

MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols

  • 构建MCP安全分类体系,覆盖协议层与主机侧17类攻击
  • 三平台测试显示所有攻击面均可成功入侵,防护措施平均失效超70%
  • 提供可扩展的评测工具包,适合安全研究与模型开发者使用

大型语言模型(LLMs)正通过模型上下文协议(MCP)这一通用开放标准连接数据源和外部工具,广泛应用于真实场景。尽管MCP提升了智能体能力,也显著扩大了攻击面并引入新安全风险。本文首次形式化定义了安全MCP及其必要规范,构建涵盖协议层与主机侧威胁的完整安全分类,识别出四大攻击面下的17种攻击类型。基于此,提出MCPSecBench——一个系统化的安全评测基准与实验平台,集成提示数据集、MCP服务器/客户端、攻击脚本、图形化测试工具及防护机制,支持在三个主流MCP平台上评估威胁。该平台模块化设计,便于研究人员自定义组件进行严格测试。对三大平台的评估表明:所有攻击面均能成功攻陷;核心漏洞普遍存在于Claude、OpenAI与Cursor中,而服务端与特定客户端攻击表现差异显著。现有防护机制平均成功率不足30%,基本无效。总体而言,MCPSecBench为MCP安全性评估提供了标准化方法,支持全协议层级的严谨测试。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly integrated into real-world applications via the Model Context Protocol (MCP), a universal open standard for connecting AI agents with data sources and external tools. While MCP enhances the capabilities of LLM-based agents, it also introduces new security risks and significantly expands their attack surface. In this paper, we present the first formalization of a secure MCP and its required specifications. Based on this foundation, we establish a comprehensive MCP security taxonomy that extends existing models by incorporating protocol-level and host-side threats, identifying 17 distinct attack types across four primary attack surfaces. Building on these specifications, we introduce MCPSecBench, a systematic security benchmark and playground that integrates prompt datasets, MCP servers, MCP clients, attack scripts, a GUI test harness, and protection mechanisms to evaluate these threats across three major MCP platforms. MCPSecBench is designed to be modular and extensible, allowing researchers to incorporate custom implementations of clients, servers, and transport protocols for rigorous assessment. Our evaluation across three major MCP platforms reveals that all attack surfaces yield successful compromises. Core vulnerabilities universally affect Claude, OpenAI, and Cursor, while server-side and specific client-side attacks exhibit considerable variability across different hosts and models. Furthermore, current protection mechanisms proved largely ineffective, achieving an average success rate of less than 30%. Overall, MCPSecBench standardizes the evaluation of MCP security and enables rigorous testing across all protocol layers.

模型安全协议评测大模型应用攻击面

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。