arXiv:2505.14590cs.CL2025-05EMNLP被引 51

提出MCIP协议,提升MCP生态中的模型安全防护能力。

MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol

  • 基于MAESTRO框架分析MCP安全缺失,设计改进协议MCIP。
  • 构建细粒度安全行为分类体系,开发评估基准与训练数据。
  • 实验证明该方法显著提升大模型在MCP场景下的风险识别能力。

随着模型上下文协议(MCP)为用户和开发者提供便捷生态,其未被充分探索的安全风险也日益凸显。其去中心化架构将客户端与服务器分离,给系统性安全分析带来独特挑战。本文提出一种新框架以增强MCP安全性。受MAESTRO框架指导,我们首先分析MCP中缺失的安全机制,并据此提出模型上下文完整性协议(MCIP),作为对MCP的优化版本以填补这些漏洞。随后,我们构建了一个细粒度的分类体系,涵盖MCP场景中观察到的多样化不安全行为。基于此分类体系,我们开发了基准测试集与训练数据,用于评估和提升大语言模型(LLMs)在识别MCP交互中安全风险方面的能力。利用所提出的基准与数据集,我们在主流大模型上进行了广泛实验。结果表明,当前大模型在MCP交互中存在明显脆弱性,而我们的方法能显著提升其安全表现。

原文摘要 · Abstract (English)

As Model Context Protocol (MCP) introduces an easy-to-use ecosystem for users and developers, it also brings underexplored safety risks. Its decentralized architecture, which separates clients and servers, poses unique challenges for systematic safety analysis. This paper proposes a novel framework to enhance MCP safety. Guided by the MAESTRO framework, we first analyze the missing safety mechanisms in MCP, and based on this analysis, we propose the Model Contextual Integrity Protocol (MCIP), a refined version of MCP that addresses these gaps. Next, we develop a fine-grained taxonomy that captures a diverse range of unsafe behaviors observed in MCP scenarios. Building on this taxonomy, we develop benchmark and training data that support the evaluation and improvement of LLMs' capabilities in identifying safety risks within MCP interactions. Leveraging the proposed benchmark and training data, we conduct extensive experiments on state-of-the-art LLMs. The results highlight LLMs' vulnerabilities in MCP interactions and demonstrate that our approach substantially improves their safety performance.

模型安全MCPLLM评估协议设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。