arXiv:2604.08805cs.CRcs.AI2026-04被引 2

构建更优强化学习环境,提升自主网络防御能力

Building Better Environments for Autonomous Cyber Defence

  • 提出接口分解框架,连接强化学习环境与真实系统
  • 总结行业最佳实践,指导防御环境开发与评估
  • 适合从事自主网络防御研究的学者与工程师

2025年11月,作者组织了一场关于强化学习(RL)在自主网络防御(ACD)中应用的研讨会。本文汇总了参会者在研讨期间及会后分享的知识,涵盖来自学术界、产业界和政府机构的专家,他们均具备丰富的RL与网络安全环境设计经验。尽管已有大量关于RL用于ACD的研究文献,但实践中积累的工程经验、领域知识和常见陷阱仍缺乏系统性整合。本文聚焦于构建更优的训练与评估环境,支持在政府与关键基础设施网络场景下训练和评估自主RL智能体。贡献包括:(1) 一套用于分解RL网络安全环境与真实系统之间接口的框架;(2) 基于研讨会核心发现,提出的当前最佳实践指南,涵盖基于强化学习的ACD环境开发与代理评估方法。

原文摘要 · Abstract (English)

In November 2025, the authors ran a workshop on the topic of what makes a good reinforcement learning (RL) environment for autonomous cyber defence (ACD). This paper details the knowledge shared by participants both during the workshop and shortly afterwards by contributing herein. The workshop participants come from academia, industry, and government, and have extensive hands-on experience designing and working with RL and cyber environments. While there is now a sizeable body of literature describing work in RL for ACD, there is nevertheless a great deal of tradecraft, domain knowledge, and common hazards which are not detailed comprehensively in a single resource. With a specific focus on building better environments to train and evaluate autonomous RL agents in network defence scenarios, including government and critical infrastructure networks, the contributions of this work are twofold: (1) a framework for decomposing the interface between RL cyber environments and real systems, and (2) guidelines on current best practice for RL-based ACD environment development and agent evaluation, based on the key findings from our workshop.

自主防御强化学习网络安全环境设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。