arXiv:2508.19278cs.CRcs.AI2025-08被引 2

扩展CybORG环境,让AI在模拟网络攻防中更接近真实操作。

Towards Production-Worthy Simulation for Autonomous Cyber Operations

  • 新增修补、隔离、解除隔离三类操作,贴近真实运维行为。
  • 改进奖励机制与特征空间,使DQN和PPO训练效果提升。
  • 适合研究网络安全RL训练的学者与安全自动化开发者。

模拟环境在自主网络攻防(ACO)中至关重要,使强化学习(RL)代理可在无仿真计算开销的情况下进行训练。这些环境需准确反映网络安全场景,并生成支持RL训练的有效信号。本研究首先在CybORG的Cage Challenge 2环境中扩展了三种新动作:修补(Patch)、隔离(Isolate)和解除隔离(Unisolate),以更好体现人类操作员在真实场景中的能力。随后提出一种代理开发设计,通过调整奖励信号与代理特征空间来提升训练性能。我们在更新后的环境中训练了DQN和PPO代理以验证改进效果。结果表明,CybORG可通过增加现实功能进行扩展,同时仍能为RL代理生成有效的训练信号。

原文摘要 · Abstract (English)

Simulated environments have proven invaluable in Autonomous Cyber Operations (ACO) where Reinforcement Learning (RL) agents can be trained without the computational overhead of emulation. These environments must accurately represent cybersecurity scenarios while producing the necessary signals to support RL training. In this study, we present a framework where we first extend CybORG's Cage Challenge 2 environment by implementing three new actions: Patch, Isolate, and Unisolate, to better represent the capabilities available to human operators in real-world settings. We then propose a design for agent development where we modify the reward signals and the agent's feature space to enhance training performance. To validate these modifications, we train DQN and PPO agents in the updated environment. Our study demonstrates that CybORG can be extended with additional realistic functionality, while maintaining its ability to generate informative training signals for RL agents.

网络安全强化学习模拟环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。