arXiv:2502.08586cs.LGcs.AI2025-02被引 56

商用大模型代理易受简单攻击,安全风险远超孤立模型。

Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks

  • 构建攻击分类体系,覆盖角色、目标、入口等维度
  • 实测主流代理系统,攻击可零基础实现且效果显著
  • 适合关注模型部署安全的研究者与开发者

当前大量机器学习安全研究聚焦于对对齐大语言模型(LLMs)的攻击,这些攻击可能窃取隐私信息或诱导模型生成有害输出。在真实部署中,大模型常作为更大智能体流程的一部分,包括记忆系统、检索模块、网络访问和API调用等。这些附加组件引入了新的漏洞,使大模型驱动的智能体比孤立的模型更容易被攻击,但相关安全研究相对匮乏。本文分析了大模型代理独有的安全与隐私漏洞,首先提出一个涵盖威胁者、目标、入口点、攻击者可观测性、攻击策略及代理流水线固有缺陷的攻击分类体系。随后,我们对流行的开源与商业代理进行了系列演示性攻击,揭示其漏洞的即时实际影响。值得注意的是,这些攻击实现极其简单,无需机器学习知识即可完成。

原文摘要 · Abstract (English)

A high volume of recent ML security literature focuses on attacks against aligned large language models (LLMs). These attacks may extract private information or coerce the model into producing harmful outputs. In real-world deployments, LLMs are often part of a larger agentic pipeline including memory systems, retrieval, web access, and API calling. Such additional components introduce vulnerabilities that make these LLM-powered agents much easier to attack than isolated LLMs, yet relatively little work focuses on the security of LLM agents. In this paper, we analyze security and privacy vulnerabilities that are unique to LLM agents. We first provide a taxonomy of attacks categorized by threat actors, objectives, entry points, attacker observability, attack strategies, and inherent vulnerabilities of agent pipelines. We then conduct a series of illustrative attacks on popular open-source and commercial agents, demonstrating the immediate practical implications of their vulnerabilities. Notably, our attacks are trivial to implement and require no understanding of machine learning.

大模型安全智能体攻击隐私泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。