arXiv:2505.23634cs.LGcs.CR2025-05被引 3

用新方法训练大模型拒绝伪造良性攻击,提升AI安全防护能力。

MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment

  • 构建MCP伪造良性攻击数据集,用于直接偏好优化训练
  • 发现不同对齐方式下模型拒绝对策能力差异显著,部分模型效果极差
  • 提出RAG-Pref新策略,结合DPO显著提升模型拒绝对抗能力

模型上下文协议(MCP)作为生成式AI代理的开放标准被广泛采用,但近期研究显示其易受基于检索的“伪造良性攻击”(FBAs)影响,导致恶意系统访问和凭证窃取,此前认为需用户主动下载受损文件。本文揭示攻击威胁远超预期——攻击者仅需在线发布恶意内容即可诱导MCP代理在不知情受害者系统上执行攻击。为强化模型对这类攻击的防御能力,我们构建了包含伪造良性与真实良性样本的MCP-FBA数据集,探索直接偏好优化(DPO)在大语言模型拒绝对话训练中的有效性。结果显示,尽管DPO能提升模型防御能力,但其效果随原始模型对齐方式差异极大,例如基于GRPO的模型拒绝对策能力极弱。为此,我们提出基于检索增强生成的偏好对齐方法(RAG-Pref),实验表明其显著提升了模型对FBAs的拒绝能力,尤其与DPO结合时,大幅增强了对基于MCP攻击的防护防线。

原文摘要 · Abstract (English)

The model context protocol (MCP) has been widely adapted as an open standard enabling the seamless integration of generative AI agents. However, recent work has shown the MCP is susceptible to retrieval-based "falsely benign" attacks (FBAs), allowing malicious system access and credential theft, but requiring that users download compromised files directly to their systems. Herein, we show that the threat model of MCP-based attacks is significantly broader than previously thought, i.e., attackers need only post malicious content online to deceive MCP agents into carrying out their attacks on unsuspecting victims' systems. To improve alignment guardrails against such attacks, we introduce a new MCP dataset of FBAs and (truly) benign samples to explore the effectiveness of direct preference optimization (DPO) for the refusal training of large language models (LLMs). While DPO improves model guardrails against such attacks, we show that the efficacy of refusal learning varies drastically depending on the model's original post-training alignment scheme--e.g., GRPO-based LLMs learn to refuse extremely poorly. Thus, to further improve FBA refusals, we introduce Retrieval Augmented Generation for Preference alignment (RAG-Pref), a novel preference alignment strategy based on RAG. We show that RAG-Pref significantly improves the ability of LLMs to refuse FBAs, particularly when combined with DPO alignment, thus drastically improving guardrails against MCP-based attacks.

AI安全偏好对齐大模型防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。