arXiv:2506.23260cs.CRcs.AI2025-06被引 107

系统梳理LLM智能体生态中的攻击威胁,给出可落地的安全框架。

From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows

  • 构建端到端威胁模型,覆盖指令注入与协议漏洞
  • 归纳30余种攻击技术,涵盖数据、模型、系统层风险
  • 适合安全研究人员和智能体系统设计者参考

由大语言模型驱动的自主智能体通过结构化函数调用接口实现实时数据获取、计算与多步编排。然而,插件、连接器及跨智能体协议的快速扩展远超安全实践,导致集成脆弱,依赖临时认证、不一致的模式和薄弱验证。本文提出一个统一的端到端威胁模型,涵盖主机-工具与智能体-智能体间的通信。系统性分类超过三十种攻击技术,包括输入操纵、模型劫持、系统与隐私攻击以及协议层漏洞。每类均提供形式化威胁定义,明确攻击者能力、目标与受影响层级。代表性案例包括Prompt-to-SQL注入及GitHub MCP服务器中的毒化智能体流攻击。分析攻击可行性,回顾现有防御措施,并讨论动态信任管理、加密溯源追踪与沙箱接口等缓解策略。该框架经专家评审并映射真实事件与公开漏洞库(如CVE、NIST NVD)验证。相比已有综述,本工作首次整合输入级漏洞与协议层缺陷,为设计安全可靠的智能体系统提供可操作指导。

原文摘要 · Abstract (English)

Autonomous AI agents powered by large language models (LLMs) with structured function-calling interfaces enable real-time data retrieval, computation, and multi-step orchestration. However, the rapid growth of plugins, connectors, and inter-agent protocols has outpaced security practices, leading to brittle integrations that rely on ad-hoc authentication, inconsistent schemas, and weak validation. This survey introduces a unified end-to-end threat model for LLM-agent ecosystems, covering host-to-tool and agent-to-agent communications. We systematically categorize more than thirty attack techniques spanning input manipulation, model compromise, system and privacy attacks, and protocol-level vulnerabilities. For each category, we provide a formal threat formulation defining attacker capabilities, objectives, and affected system layers. Representative examples include Prompt-to-SQL injections and the Toxic Agent Flow exploit in GitHub MCP servers. We analyze attack feasibility, review existing defenses, and discuss mitigation strategies such as dynamic trust management, cryptographic provenance tracking, and sandboxed agent interfaces. The framework is validated through expert review and cross-mapping with real-world incidents and public vulnerability repositories, including CVE and NIST NVD. Compared to prior surveys, this work presents the first integrated taxonomy bridging input-level exploits and protocol-layer vulnerabilities in LLM-agent ecosystems, offering actionable guidance for designing secure and resilient agentic AI systems.

智能体安全威胁建模协议漏洞LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。