arXiv:2608.15012cs.CRcs.AI2026-08

构建自驱式攻防共进化系统,实现安全能力的持续迭代。

SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system

论文配图:SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system
图 1 · 摘自论文原文
  • 攻防双智能体通过对抗推动彼此演化,形成闭环。
  • 攻击成功率提升25%以上,防御精度达百至千倍提升。
  • 适合安全研究、红蓝对抗及AI安全评估场景。

大语言模型的快速发展加剧了网络安全中攻防不对称性:攻击趋向自主执行,而防御仍高度依赖人工。尽管已有大量工作涉及网络靶场、AI驱动攻击与防御,这种不对称依然存在。我们发现其根源在于攻防演进在三个层面均停滞不前。为此,提出以共进化为核心理念,让攻击与防御的AI智能体在对抗中自主且安全地相互驱动演化。基于此,构建了 exttt{SysEvolve} 系统,包含 exttt{SysField}、 exttt{SysSpear} 与 exttt{SysArmor} 三组件: exttt{SysField} 构建真实多主机靶场; exttt{SysSpear} 生成高效安全的攻击方案; exttt{SysArmor} 实现实时可解释的防御。三者构成自驱式对抗循环,恢复了三层演进能力。评估显示, exttt{SysField} 在2.1%开销下实现零损失数据采集,并将257个CVE整合为1,148个靶场; exttt{SysSpear} 攻击成功率超过基线模型25%; exttt{SysArmor} 精度较以往系统提升10–1000倍,并在华为与深信服生产环境检测到真实APT攻击。评估还揭示三点:多步组合与更大拓扑暴露了单步评估掩盖的能力短板;瓶颈出现在初始访问后的横向移动阶段;LLM智能体易受环境干扰——部署诱饵端点后,超时次数增至三倍,下游任务完全失败,尽管初始访问成功率不变。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while defense remains predominantly human-intensive. Despite substantial prior work across cyber ranges, AI-driven attack, and AI-driven defense, this asymmetry persists. We trace it to a deeper root cause, that evolution itself has stalled on both sides at three layers. To overcome this, we propose co-evolution as the integrating insight, where attack and defense AI agents autonomously and safely drive each other's evolution through adversarial confrontation. Based on this insight, we present \sysevolve, comprising three co-designed components, \sysfield, \sysspear, and \sysarmor. \sysfield constructs realistic multi-host ranges. \sysspear generates efficient, safe attack schemes. \sysarmor performs real-time, interpretable defense. Together they form a self-driven adversarial loop restoring evolution at all three layers. In evaluation, \sysfield achieves zero-loss collection at 2.1\% overhead and orchestrates 257 CVEs into 1,148 ranges, \sysspear improves attack success by over 25\% over baseline LLMs, and \sysarmor achieves 10--1000$\times$ greater precision than prior systems and detects real APT attacks in production at Huawei and Sangfor. Our evaluation also reveals three findings about LLM agent capabilities. First, multi-step composition and larger topologies expose agent capability gaps hidden by single-step evaluations. Second, the bottleneck lies after initial access in post-compromise state utilization. Third, LLM agents are susceptible to environmental interference. When decoy endpoints are deployed in the range, agent timeouts triple and downstream completion disappears despite the success rates of initial accesses are unchanged.

攻防共进化LLM安全AI防御红蓝对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。