提出混合分析框架MTGuard,防范LLM代理滥用MCP工具的安全风险。
Hybrid Analysis for Secure MCP Tool Use in LLM Agents

- 结合静态与动态分析,按生命周期检测MCP工具调用行为。
- 在多种代理上有效阻止恶意工具使用,同时保持正常任务性能。
- 适合关注LLM安全、工具调用防护的研究者和开发者。
大语言模型(LLM)代理的快速发展使其广泛应用于各类实际任务。为标准化LLM代理与外部环境的交互,模型上下文协议(MCP)工具已成为事实标准,并被广泛集成到系统中。然而,MCP工具的使用也引入了新的安全风险,即LLM代理可能被诱导执行恶意或未经授权的操作。尽管已有研究提出针对LLM代理工具使用的防御方法,但多数依赖静态分析(检查提示词和生成输出),限制了防御效果与鲁棒性。为此,我们提出MTGuard,一种基于混合分析的防御框架,通过生命周期感知的静态-动态协同分析,保护LLM代理中MCP工具的使用安全。大量实验表明,MTGuard能有效缓解多种类型的有害工具使用,同时在良性用户任务上保持良好性能。
原文摘要 · Abstract (English)
The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize interactions between LLM agents and external environments, Model Context Protocol (MCP) tools have emerged as a de facto standard and have been widely integrated into these systems. However, the use of MCP tools also introduces new safety risks, as LLM agents can be induced to perform malicious or unauthorized actions. Although prior work has proposed defenses for securing tool use in LLM agents, most methods rely on static analysis, i.e., inspecting prompts and generated outputs, which limits the defense effectiveness and robustness. To address these limitations, we propose MTGuard, a hybrid analysis-based defense framework designed to safeguard the use of MCP tools in LLM agents by leveraging lifecycle-aware static-dynamic co-analysis. Extensive evaluation demonstrates that MTGuard effectively mitigates multiple categories of harmful tool use across different LLM agents while maintaining performance on benign user tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。