arXiv:2606.24496cs.CRcs.AI2026-06

首次深度分析进攻性智能体安全漏洞,揭示其可被远程窃取密钥并突破沙箱。

Red-Teaming the Agentic Red-Team

  • 构建完整攻击链,从大模型操控到沙箱逃逸
  • 多数工具存在共性缺陷,可在沙箱内完全攻陷主机
  • 提出可复用的架构设计原则,提升智能体安全性

使用智能体执行进攻性安全操作已从理论变为现实能力。然而,尽管社区持续提升智能体能力,对其安全性的评估却相对不足。本文首次对最广泛使用的进攻性智能体系统进行深入安全分析,发现多数工具存在共性设计缺陷,使主动攻击者能窃取API密钥、建立持久后门,并在沙箱容器内完全控制操作者机器。为支持分析,我们构建了完整的网络杀伤链,涵盖从初始大模型操控到横向移动、持久化、护栏绕过及沙箱逃逸的全过程。基于此,我们提出一种稳健的智能体进攻性工具架构,并制定可广泛适用的设计原则,从架构层面防御已披露的攻击路径。

原文摘要 · Abstract (English)

The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, while the community has focused on creating more and more capable agents, less attention has been allocated to assessing the security of those systems. In this work, we present the first in-depth security analysis of the most widely used agentic systems for offensive security operations. We show that most of these tools share common design flaws that enable an active adversary to exfiltrate API keys, establish persistent footholds, and fully compromise the operator's machine, even when the agent operates inside a sandboxed container. To support our analysis, we introduce a full cyber kill chain for such agentic systems, capturing the progression from initial LLM manipulation to lateral movement, persistence, guardrail bypass, and sandbox escape. Building on our security analysis, we derive a robust architecture for agentic offensive-security tools and propose actionable, broadly applicable design principles that mitigate the disclosed attack paths at the architectural level.

智能体安全攻击链沙箱逃逸

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。