攻击者通过伪造MCP服务器诱导大模型优先使用其服务,谋取经济利益。
MPMA: Preference Manipulation Attack Against Model Context Protocol
- 设计直接操纵工具名称和描述的攻击方法,实现偏好操控。
- 提出基于遗传算法的隐蔽攻击方案,平衡效果与隐匿性。
- 揭示MCP开放生态中的安全漏洞,适合关注AI安全的研究者阅读。
模型上下文协议(MCP)为大语言模型(LLMs)访问外部数据和工具提供了标准化接口,推动了工具选择范式的变革,并加速了LLM智能体工具生态的发展。然而,随着MCP的广泛应用,第三方定制的MCP服务器暴露了潜在的安全风险。本文首次提出一种新型安全威胁——MCP偏好操纵攻击(MPMA):攻击者部署自定义MCP服务器,操纵大模型优先选择自身服务,从而获取经济收益,如付费服务收入或免费服务器带来的广告收入。为实现该攻击,我们首先设计了直接偏好操纵攻击(DPMA),通过在工具名称和描述中插入操纵性词汇取得显著效果,但此类修改易被用户察觉,缺乏隐蔽性。为此,我们进一步提出基于遗传算法的广告偏好操纵攻击(GAPMA),采用四种常见策略初始化描述,并引入遗传算法优化以增强隐蔽性。实验结果表明,GAPMA在保持高攻击效果的同时具备更强的隐蔽性。本研究揭示了开放生态系统中MCP的关键安全漏洞,凸显了构建鲁棒防御机制以保障MCP生态公平性的紧迫性。
原文摘要 · Abstract (English)
Model Context Protocol (MCP) standardizes interface mapping for large language models (LLMs) to access external data and tools, which revolutionizes the paradigm of tool selection and facilitates the rapid expansion of the LLM agent tool ecosystem. However, as the MCP is increasingly adopted, third-party customized versions of the MCP server expose potential security vulnerabilities. In this paper, we first introduce a novel security threat, which we term the MCP Preference Manipulation Attack (MPMA). An attacker deploys a customized MCP server to manipulate LLMs, causing them to prioritize it over other competing MCP servers. This can result in economic benefits for attackers, such as revenue from paid MCP services or advertising income generated from free servers. To achieve MPMA, we first design a Direct Preference Manipulation Attack (DPMA) that achieves significant effectiveness by inserting the manipulative word and phrases into the tool name and description. However, such a direct modification is obvious to users and lacks stealthiness. To address these limitations, we further propose Genetic-based Advertising Preference Manipulation Attack (GAPMA). GAPMA employs four commonly used strategies to initialize descriptions and integrates a Genetic Algorithm (GA) to enhance stealthiness. The experiment results demonstrate that GAPMA balances high effectiveness and stealthiness. Our study reveals a critical vulnerability of the MCP in open ecosystems, highlighting an urgent need for robust defense mechanisms to ensure the fairness of the MCP ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。