arXiv:2607.26314cs.CRcs.AI2026-07

测试自主安全工具的隐蔽性,发现多数模型完成任务却暴露行踪。

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

  • 构建六维隐蔽性评测基准,模拟真实攻防场景
  • 14个任务中仅54%模型既完成任务又保持隐蔽
  • 适合研究自主攻防系统或安全监控的开发者

隐蔽性是达成目标而不暴露自身存在、能力或情报的关键,区分了高阶安全研究人员与可被检测的攻击者。随着自主代理承担越来越多的进攻性任务,它们是否继承了相应的隐蔽技巧?我们提出StealthBench,一个衡量自主进攻型安全代理在六个操作安全(OPSEC)维度上的隐蔽性基准。从真实的漏洞赏金和红队轨迹中提取11个经人工验证的OPSEC事件,扩展为14个容器化任务场景。尽管代理发现了真实漏洞,仍出现系统性隐蔽失败:将凭证嵌入公开上传、删除生产资源以证明访问权限、强制添加无关用户以展示竞争条件。我们采用三模型大语言模型判官组进行多数投票评估,测量安全成功率(任务完成且隐蔽)、Stealth@Solve(成功解题中的隐蔽质量)以及鲁莽解决率(完成但暴露)。结果表明,无一模型安全成功率超过54%,证实了跨模型家族的系统性OPSEC缺陷。我们公开发布StealthBench,支持隐蔽感知代理开发及自主进攻部署的自动化OPSEC监控。交互式排行榜、评估工具包和数据集可在https://stealthbench.com获取。

原文摘要 · Abstract (English)

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions. We extract 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios, where agents, despite finding real vulnerabilities, committed stealth failures inconsistent with standard operational tradecraft: embedding credentials in public uploads, deleting production resources to prove access, force-adding uninvolved users to demonstrate a race condition. We evaluate agent trajectories using a 3-model large language model (LLM) judge panel with majority-vote aggregation, measuring safe success rate (solved and stealthy), Stealth@Solve (tradecraft quality among successful solves), and reckless solve rate (solved but cover blown). Our results show that no model exceeds 54% safe success rate (the compound metric requiring both task completion and stealth), confirming that OPSEC failures are systematic across model families. We release StealthBench as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments. The interactive leaderboard, evaluation harness, and dataset are available at https://stealthbench.com.

自主攻防隐蔽性评测安全基准大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。