在网络安全攻防中,技能包反而降低模型表现,因环境反馈已足够提供纠错信号。
When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity

- 用四种文档丰富度测试技能包效果,发现越完整反而越差
- 技能包使任务通过率仅提升8.9个百分点,且不显著(p=0.71)
- 适合研究强化学习与环境反馈机制的开发者参考
Agent Skills 是推理时加载到 LLM 代理中的结构化过程知识,被广泛报道可使任务通过率平均提高 16.2~个百分点。然而,同一基准测试显示,84 个任务中有 16 个引入技能后通过率下降。我们重新分析一项近期发布的 180 次受控实验,研究基于 MCP 的自主攻防(CTF)代理在四种文档条件(591、12865、17253 和 36001 字符)下的表现,对应无技能、经验型技能、精选技能和全面技能的消融设置。在攻击性网络安全领域,现有技能包并未带来明显收益:无技能与全技能条件下通过率差距仅为 8.9 个百分点(p = 0.71, χ²;p = 0.25, Cochran–Armitage 趋势检验;六组配对 Cohen's h 值中五组低于 0.2 的小效应阈值)。我们提出,缺失变量是环境反馈带宽——当工具层返回严格、结构验证、低延迟的观测结果时,环境本身已能提供过程修正信号,使技能包变得冗余甚至有害。本文提出可证伪假说,并讨论复合智能系统的设计启示,将公开重分析流程以支持复现。
原文摘要 · Abstract (English)
Agent Skills, structured packages of procedural knowledge loaded into an LLM agent at inference time, are widely reported to improve task pass rates by an average of 16.2~percentage points across diverse domains. Yet the same benchmarks show wide variance, with 16 of 84 tasks suffering negative deltas when Skills are introduced. The community has not yet articulated a clean mechanism for \emph{when} Skills help and when they are merely redundant overhead. We re-analyze a recently published 180-run controlled study of an MCP-grounded autonomous Capture-the-Flag (CTF) agent under four documentation conditions of increasing richness (591, 12865, 17253, and 36001 tokens) and show that these conditions correspond almost exactly to a No-Skills, Experiential-Skills, Curated-Skills, and Comprehensive-Skills ablation. In offensive cybersecurity, a domain not deeply covered by existing Skills benchmarks, the marginal benefit of Skills collapses. The spread between the no-Skills and full-Skills conditions is only 8.9~pp ($p = 0.71$, $χ^2$; $p = 0.25$, Cochran--Armitage trend test; five of six pairwise Cohen's $h$ values fall below the $0.2$ small-effect threshold). We argue that the missing variable is \emph{environment-feedback bandwidth}. When an agent's tool layer returns strict, schema-validated, low-latency observations, the environment itself supplies the procedural correction signal that Skills are normally needed to provide. As a result, the marginal benefit of curated Skills diminishes substantially, and, in some cases (e.g., our timing side-channel setting), actively degrades performance. We articulate a falsifiable hypothesis, sketch its design implications for compound AI systems, and will release the reanalysis pipeline to support replication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。