arXiv:2607.26998cs.CRcs.CL2026-07被引 2

动态构建假环境诱骗智能渗透代理,延缓攻击并使其误判目标。

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

论文配图:AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents
图 1 · 摘自论文原文
  • 根据代理行为历史动态生成欺骗性假设施
  • 吸收46.8%工具调用,55.9%攻击动作滞留假环境
  • 90%完成报告基于假证据,3次尝试无真实目标被攻破

大型语言模型(LLM)代理通过观察-动作循环自动化渗透测试,依赖工具返回的观察结果选择行动。这使得防御方可通过注入虚假观察误导其决策。然而现有防御高度依赖攻击前预设的静态、孤立伪造物,高级代理能逐步识别并绕过这些伪造物,最终重新聚焦于真实目标。为此,我们提出AgentSnare,一种轨迹自适应欺骗系统,可动态展开假环境以持续引导渗透代理远离真实目标。AgentSnare采用基于交互历史与假环境状态的伪物构造策略模型,生成候选伪造物;随后验证并逐步将有效伪造物整合至语义一致的假环境中,从而延迟攻击(吸收46.8%工具调用)、引导攻击路径(55.9%后进入动作滞留假环境),并通过诱导基于假证据的完成报告实现消解。在15个CVE-Bench Web应用和3种攻击者模型下,所有45个攻击者-CVE组合中,pass@3下均未成功利用真实目标。

原文摘要 · Abstract (English)

Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process. However, existing defenses rely heavily on static, isolated artifacts planted in the environment prior to an attack. Advanced agents can progressively recognize and bypass these artifacts, ultimately refocusing their exploitation attempts on the real target. To address this issue, we introduce AgentSnare, a trajectory-adaptive deception system that dynamically unfolds a decoy environment to continually steer the penetration agent away from the real target. Specifically, AgentSnare employs an artifact-construction policy model that constructs candidate artifacts conditioned on the agent's interaction history and decoy state. AgentSnare then validates these candidates and incrementally incorporates valid artifacts into a factually consistent decoy environment, thereby delaying the attack by absorbing its tool calls, diverting its post-entry trajectory within the decoy, and defusing it by inducing completion reports grounded in decoy evidence. Across 15 CVE-Bench web applications and three attacker models, AgentSnare absorbs 46.8% of the agent's tool calls in the decoy and retains 55.9% of post-entry actions there, while 90.0% of completion attempts are grounded in decoy evidence; across all 45 attacker-CVE pairs, no real target is successfully exploited at pass@3.

AI安全欺骗防御渗透测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。