arXiv:2603.11214cs.AIcs.LG2026-03被引 11

测试顶尖AI在复杂网络攻击中的自主能力,发现算力越大越强,模型迭代进步明显。

Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios

  • 用真实攻防场景测试AI链式行动能力,覆盖32步企业网络攻击和7步工控系统攻击
  • 算力翻10倍性能提升59%,模型每代都比前代更强,最先进模型完成22步(共32步)
  • 适合关注AI安全风险、攻防对抗与模型演进的研究者和安全从业者

我们评估了前沿人工智能模型在两个定制化网络安全环境中的自主攻击能力——一个包含32个步骤的企业网络攻击场景和一个包含7个步骤的工业控制系统攻击场景,这些场景需要跨多个异构任务的长期动作序列。通过比较18个月(2024年8月至2026年2月)间发布的七种模型在不同推理计算预算下的表现,观察到两个趋势:第一,模型性能随推理时计算量呈对数线性增长,未见饱和;从1000万到1亿次令牌推理,性能最高提升59%,无需操作者具备特殊技术能力;第二,每一代新模型在固定令牌预算下均优于前代:在企业网络场景中,1000万令牌下平均完成步骤数从GPT-4o(2024年8月)的1.7升至Opus 4.6(2026年2月)的9.8;最佳单次运行完成22/32步,相当于人类专家约14小时工作量的6小时。在工业控制系统场景中,性能仍有限,但最新模型首次能稳定完成部分步骤,平均完成1.2–1.4/7步(最高3步)。

原文摘要 · Abstract (English)

We evaluate the autonomous cyber-attack capabilities of frontier AI models on two purpose-built cyber ranges-a 32-step corporate network attack and a 7-step industrial control system attack-that require chaining heterogeneous capabilities across extended action sequences. By comparing seven models released over an eighteen-month period (August 2024 to February 2026) at varying inference-time compute budgets, we observe two capability trends. First, model performance scales log-linearly with inference-time compute, with no observed plateau-increasing from 10M to 100M tokens yields gains of up to 59%, requiring no specific technical sophistication from the operator. Second, each successive model generation outperforms its predecessor at fixed token budgets: on the corporate network range, average steps completed at 10M tokens rose from 1.7 (GPT-4o, August 2024) to 9.8 (Opus 4.6, February 2026). The best single run completed 22 of 32 steps, corresponding to roughly 6 of the estimated 14 hours a human expert would need. On the industrial control system range, performance remains limited, though the most recent models are the first to reliably complete steps, averaging 1.2-1.4 of 7 (max 3).

AI安全攻防测试智能代理模型演进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。