arXiv:2608.21423cs.CLcs.AI2026-08

用大模型代理做渗透测试,揭示了常见失败原因和设计规律。

Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing

  • 构建四维集成摩擦指数,区分一次性工程与持续成本
  • 发现长会话易丢失证据,短子代理能延长可用周期
  • 指出不能靠提示词控制权限,需显式机制保障安全

智能安全利用大语言模型(LLM)代理规划、调度并解释安全工具。随着系统从演示走向部署,从业者反复遭遇相同操作失败。我们通过实测十种常用静态、动态、云、编排及AI红队工具,系统化总结这些失败模式。提出四维集成摩擦指数,分离一次性工程成本与持续的组织、法律和维护成本。建模为随机LLM策略被确定性中介包裹,发现长期会话随阶段数丢失留存证据,而短期子代理可依据原始证据与摘要的压缩比扩展可用时长。两阶段判断级联虽提升评分器似然比,但当评分错误相关时收益甚微。将不可评估结果视为攻击失败会误导下游测量,偏向回避和严重响应。将规划者-工作者路由建模为背包问题,推导出重尾工具的闭式执行上限:eta* = alpha v/c。最后证明,范围与预算约束无法委托给系统提示词——提示词无法限制实际执行内容。Inspectra平台作为实现范例,标注了已交付、部分实现或计划中的机制,包括未成功实现的方案。

原文摘要 · Abstract (English)

Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures. We systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pipelines. We introduce a four-dimensional Integration Friction Index that separates one-time engineering cost from recurring organisational, legal, and maintenance cost. We then derive quantitative regularities that explain recurring failure modes. Modelling an agentic security system as stochastic LLM policies wrapped by a deterministic mediator, we show that long-lived sessions lose resident evidence with phase count, while short-lived sub-agents extend the usable horizon according to the compression ratio between raw evidence and its summary. We show that a two-stage verdict cascade multiplies scorer likelihood ratios, but provides little benefit when scorer errors correlate. We show that treating unevaluable outcomes as attack failures biases downstream measurements toward evasive and severe responses. We formulate planner-versus-worker model routing as a knapsack problem and derive a closed-form execution cap for heavy-tailed tools, eta* = alpha v/c. Finally, we show why scope and budget enforcement cannot be delegated to system prompts: prompts do not constrain what actually executes. Inspectra, our implemented platform, serves as a worked instantiation, with mechanisms labelled shipped, partial, or planned, including those that did not work.

智能安全大模型代理渗透测试系统设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。