arXiv:2607.11288cs.CRcs.AI2026-07

Mako能自动发现并修复漏洞,实现无人值守的网络攻防测试。

Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation

论文配图:Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation
图 1 · 摘自论文原文
  • 将漏洞利用能力视为可动态升级的内核,实时自我进化。
  • 在104个靶机上全部触发唯一指纹标志,验证结果不可伪造。
  • 适合安全研究者和自动化渗透测试平台开发者参考。

我们提出自演化代理操作系统(SE-AOS):一种新型人工智能代理,将漏洞利用能力视为可运行时更新、版本化的内核。它能观察自身失败,合成新能力,在真实目标上验证后热加载回自身。Mako是首个用于安全研究的SE-AOS实例,也是LaunchSafe平台的核心引擎。在公开的XBOW基准测试中,针对104个容器化CTF风格的网页应用(涵盖26类漏洞,分三个难度层级),Mako实现了全量覆盖:每个目标均生成唯一的、加密新鲜的构建级标志,验证机制杜绝了伪造或记忆结果。核心发现是:一旦能力存在且可被发现,攻击难度即刻坍缩;稀缺的是能力本身,而非推理。该架构与形式化方法使这一规律转化为自增强系统。Mako还具备受控的自我演化循环,可在性能不下降时提议、沙箱测试并提交对自身代理与规则的改进。出于双重用途风险考虑,我们刻意未公开实际攻击结果、有效载荷、漏洞链及工具源码——一个将全谱网络攻防简化为可重复、机器级速度流水线的系统,属于敏感研究。我们发布科学,但不提供武器。

原文摘要 · Abstract (English)

We introduce the Self-Evolving Agentic Operating System (SE-AOS): a new class of AI agent that treats exploit capability as a mutable, versioned kernel it extends at runtime, observing its own failures, synthesising new capabilities, proving them against a live target, and hot-loading them back into itself. Mako is the first SE-AOS instance for security research and the autonomous web exploitation engine developed within LaunchSafe. LaunchSafe builds autonomous security agents for continuous offensive testing and agent-driven security research; Mako is the core engine behind that platform. On the public XBOW validation-benchmarks, 104 containerised, CTF-style web applications spanning 26 vulnerability classes across three difficulty tiers, Mako achieves full-suite coverage: it drives every one of the 104 targets to emit a cryptographically fresh, per-build flag, under a verification regime that makes fabricated or memorised results impossible. Our central result is a law of autonomous exploitation: once a capability exists and is discoverable, difficulty collapses; capability, not reasoning, is what is scarce, together with an architecture and formalism that turn that law into a self-improving system. Mako further runs a gated self-evolution loop that proposes, sandboxes, and commits improvements to its own agents and rules when fitness does not regress. We deliberately withhold the operational results, payloads, exploit chains, and tool source, because a system that reduces full-spectrum web exploitation to a repeatable, machine-speed pipeline is dual-use research of concern. We publish the science; we withhold the weapon.

自主攻防安全代理自演化系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。