通过追踪操作轨迹,识别大模型在网络安全竞赛中是否真攻破了漏洞。
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation

- 构建可追溯的审计框架,拆解智能体行为为渗透阶段与技术类型。
- 仅62%-87%的夺旗成功由真实攻击验证,多数为捷径获取。
- 适合关注模型真实攻防能力评估的研究者与安全测试人员。
捕获旗帜(CTF)基准广泛用于评估自主语言模型智能体的进攻性安全能力。现有评估依赖浅层二元判断或聚合分数,忽视智能体获取旗帜的实际路径,导致真实攻破、直接暴露、记忆复现、外部查询、猜测及无依据宣称被混为一谈,可能高估模型的网络安全能力。我们提出CTF-ABACUS,一种基于轨迹的智能体审计框架,将每次运行重构为证据支撑的求解画像。通过将智能体行为分解为渗透测试阶段与类别化技术,该框架可定位实际攻击发生位置、旗帜首次出现时刻,并判断回收的旗帜是否由可验证行为支持。跨智能体聚合求解画像生成挑战特征,揭示成功是通过预期攻击实现,还是通过捷径路径达成。我们在240个挑战上对六种前沿及开源模型的1,435次尝试应用CTF-ABACUS,生成2,870份求解画像,采用两种评判视角。经轨迹验证的攻击仅占各基准中62%-87%的夺旗成果,而捷径恢复的路径显著更短。这一发现推动CTF评估从统计夺旗数转向验证真实攻破行为,为设计能更好隔离进攻能力的基准提供基础。
原文摘要 · Abstract (English)
Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Consequently actual exploitation is conflated with direct flag exposure, memorized recall, external lookup, guessing, and unsupported claims, potentially overstating the agent's cybersecurity capability. We introduce CTF-ABACUS, a trace-based agent auditing framework that reconstructs each run as an evidence-grounded solve profile. By decomposing agent actions into penetration-testing phases and categorical techniques, it identifies where exploitation occurs, where the flag first appears, and whether the recovered flag is supported by demonstrated behavior. Aggregating solve profiles across agents yields challenge signatures that reveal whether success was achieved via the intended exploit or via shortcut pathways. We apply CTF-ABACUS to 1,435 CTF attempts by six frontier and open-source models on 240 challenges, yielding 2,870 solve profiles under two judge lenses. Trace-verified exploits account for only 62-87% of recovered flags across benchmarks, while shortcut recoveries follow substantially shallower trajectories. These findings shift CTF evaluation from counting recovered flags to verifying demonstrated exploitation and provide a basis for designing benchmarks that better isolate the offensive capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。