arXiv:2512.15163cs.CLcs.AI2025-12被引 33

首个基于真实MCP服务器的安全评估基准,揭示大模型在多工具协作中的安全隐患。

MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers

  • 构建真实MCP服务器环境,支持跨服务器多轮交互评估
  • 发现所有主流大模型均易受20类攻击,存在安全与性能权衡
  • 适合关注AI代理安全、工具调用风险的研究者与开发者

大型语言模型正演变为能推理、规划并操作外部工具的智能体。模型上下文协议(MCP)为此提供标准化接口,连接异构工具与服务。然而,MCP的开放性与多服务器工作流引入了新安全风险,现有基准无法覆盖此类真实场景。本文提出MCP-SafetyBench,基于真实MCP服务器,支持五个领域(浏览器自动化、金融分析、位置导航、仓库管理、网页搜索)的多轮评估。该基准包含20种涵盖服务器、主机和用户端的攻击类型,并设计需多步推理与跨服务器协调的任务,在不确定性下进行测试。通过该基准系统评估主流开源与闭源大模型,发现所有模型均存在漏洞,且安全与性能间存在显著权衡。研究强调亟需强化防御机制,并确立MCP-SafetyBench作为诊断与缓解真实部署中安全风险的基础。

原文摘要 · Abstract (English)

Large language models (LLMs) are evolving into agentic systems that reason, plan, and operate external tools. The Model Context Protocol (MCP) is a key enabler of this transition, offering a standardized interface for connecting LLMs with heterogeneous tools and services. Yet MCP's openness and multi-server workflows introduce new safety risks that existing benchmarks fail to capture, as they focus on isolated attacks or lack real-world coverage. We present MCP-SafetyBench, a comprehensive benchmark built on real MCP servers that supports realistic multi-turn evaluation across five domains: browser automation, financial analysis, location navigation, repository management, and web search. It incorporates a unified taxonomy of 20 MCP attack types spanning server, host, and user sides, and includes tasks requiring multi-step reasoning and cross-server coordination under uncertainty. Using MCP-SafetyBench, we systematically evaluate leading open- and closed-source LLMs, revealing that all models remain vulnerable to MCP attacks, with a notable safety-utility trade-off. Our results highlight the urgent need for stronger defenses and establish MCP-SafetyBench as a foundation for diagnosing and mitigating safety risks in real-world MCP deployments.

大模型安全工具调用MCP协议智能体评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。