arXiv:2508.06418cs.CL2025-08被引 3

提出新方法量化MCP中对话漂移,防范外部工具引发的攻击

Quantifying Conversation Drift in MCP via Latent Polytope

  • 用潜在多面体空间建模对话轨迹,检测异常偏移
  • 在三款大模型上实现超过0.915的AUROC得分
  • 适合关注LLM安全与工具集成风险的研究者

模型上下文协议(MCP)通过集成外部工具提升大语言模型(LLMs)性能,实现动态数据聚合以增强任务执行。然而,其非隔离的执行环境带来严重安全与隐私风险。恶意内容可引发工具污染或间接提示注入,导致对话劫持、信息误导或数据泄露。现有防御措施如规则过滤或基于LLM的检测,因依赖静态特征、计算效率低且无法量化劫持程度而效果有限。为此,本文提出SecMCP框架,通过在潜在多面体空间中建模LLM激活向量,检测由外部知识诱导的对话轨迹偏移,实现对劫持、误导和数据外泄的主动预警。我们在Llama3、Vicuna、Mistral三款先进模型上,基于MS MARCO、HotpotQA、FinQA数据集进行评估,结果表明系统在保持可用性的同时,检测性能达到AUROC > 0.915。本工作贡献包括对MCP安全威胁的系统分类、基于潜在多面体的对话漂移量化新方法,以及实证验证了SecMCP的有效性。

原文摘要 · Abstract (English)

The Model Context Protocol (MCP) enhances large language models (LLMs) by integrating external tools, enabling dynamic aggregation of real-time data to improve task execution. However, its non-isolated execution context introduces critical security and privacy risks. In particular, adversarially crafted content can induce tool poisoning or indirect prompt injection, leading to conversation hijacking, misinformation propagation, or data exfiltration. Existing defenses, such as rule-based filters or LLM-driven detection, remain inadequate due to their reliance on static signatures, computational inefficiency, and inability to quantify conversational hijacking. To address these limitations, we propose SecMCP, a secure framework that detects and quantifies conversation drift, deviations in latent space trajectories induced by adversarial external knowledge. By modeling LLM activation vectors within a latent polytope space, SecMCP identifies anomalous shifts in conversational dynamics, enabling proactive detection of hijacking, misleading, and data exfiltration. We evaluate SecMCP on three state-of-the-art LLMs (Llama3, Vicuna, Mistral) across benchmark datasets (MS MARCO, HotpotQA, FinQA), demonstrating robust detection with AUROC scores exceeding 0.915 while maintaining system usability. Our contributions include a systematic categorization of MCP security threats, a novel latent polytope-based methodology for quantifying conversation drift, and empirical validation of SecMCP's efficacy.

大模型安全对话劫持潜在空间MCP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。