arXiv:2602.20720cs.CRcs.AI2026-02被引 11

提出自适应工具攻击框架,提升对智能大模型的隐蔽渗透能力。

AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs

  • 动态选择隐蔽工具并生成可迁移的对抗性提示
  • 攻击成功率提升2.13倍,系统可用性下降1.78倍
  • 适用于评估先进防御机制下的安全风险

外部数据服务(如模型上下文协议,MCP)的集成使基于大语言模型的智能体在复杂任务执行中日益强大。然而,这一进展也引入了关键安全漏洞,尤其是间接提示注入(IPI)攻击。现有攻击方法受限于静态模式,且仅在简单语言模型上评估,难以应对现代AI代理的快速演进。我们提出AdapTools,一种新型自适应IPI攻击框架,通过选择更隐蔽的攻击工具并生成自适应攻击提示,构建严格的安全部署评估环境。该方法包含两个核心组件:(1) 自适应攻击策略构建,开发可迁移的对抗性策略用于提示优化;(2) 攻击增强,识别能绕过任务相关性防御的隐蔽工具。全面实验表明,AdapTools在攻击成功率上提升2.13倍,同时使系统效用下降1.78倍。值得注意的是,该框架在面对最先进防御机制时仍保持有效性。本研究深化了对IPI攻击的理解,为未来研究提供了重要参考。

原文摘要 · Abstract (English)

The integration of external data services (e.g., Model Context Protocol, MCP) has made large language model-based agents increasingly powerful for complex task execution. However, this advancement introduces critical security vulnerabilities, particularly indirect prompt injection (IPI) attacks. Existing attack methods are limited by their reliance on static patterns and evaluation on simple language models, failing to address the fast-evolving nature of modern AI agents. We introduce AdapTools, a novel adaptive IPI attack framework that selects stealthier attack tools and generates adaptive attack prompts to create a rigorous security evaluation environment. Our approach comprises two key components: (1) Adaptive Attack Strategy Construction, which develops transferable adversarial strategies for prompt optimization, and (2) Attack Enhancement, which identifies stealthy tools capable of circumventing task-relevance defenses. Comprehensive experimental evaluation shows that AdapTools achieves a 2.13 times improvement in attack success rate while degrading system utility by a factor of 1.78. Notably, the framework maintains its effectiveness even against state-of-the-art defense mechanisms. Our method advances the understanding of IPI attacks and provides a useful reference for future research.

提示注入智能体安全自适应攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。