自动生成恶意工具,测试大模型代理在标准协议下的安全漏洞
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- 设计自动化框架,生成能伪装的恶意MCP工具
- 可成功操控主流大模型代理行为且躲避现有检测
- 适合关注AI安全与对抗攻击的研究者和开发者
大型语言模型(LLMs)的强大能力使其在多个领域广泛应用。为标准化大模型代理与其环境的交互,模型上下文协议(MCP)工具已成为事实标准,并被广泛集成到这些代理中。然而,引入MCP工具也带来了工具中毒攻击的风险,可能操纵大模型代理的行为。尽管已有研究揭示了此类漏洞,但其红队测试方法仍多处于概念验证阶段,如何在MCP工具中毒范式下实现大模型代理的自动化、系统化红队测试仍是开放问题。为此,我们提出AutoMalTool——一种通过生成恶意MCP工具来实现大模型代理自动化红队测试的框架。大规模评估表明,AutoMalTool能有效生成可操控主流大模型代理行为且规避当前检测机制的恶意工具,从而揭示了这些代理中的新安全风险。
原文摘要 · Abstract (English)
The remarkable capability of large language models (LLMs) has led to the wide application of LLM-based agents in various domains. To standardize interactions between LLM-based agents and their environments, model context protocol (MCP) tools have become the de facto standard and are now widely integrated into these agents. However, the incorporation of MCP tools introduces the risk of tool poisoning attacks, which can manipulate the behavior of LLM-based agents. Although previous studies have identified such vulnerabilities, their red teaming approaches have largely remained at the proof-of-concept stage, leaving the automatic and systematic red teaming of LLM-based agents under the MCP tool poisoning paradigm an open question. To bridge this gap, we propose AutoMalTool, an automated red teaming framework for LLM-based agents by generating malicious MCP tools. Our extensive evaluation shows that AutoMalTool effectively generates malicious MCP tools capable of manipulating the behavior of mainstream LLM-based agents while evading current detection mechanisms, thereby revealing new security risks in these agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。