提出可泛化防御不同网络结构的智能体,无需重新训练。
Towards a Generalisable Cyber Defence Agent for Real-World Computer Networks
- 用异构图神经网络生成固定大小的网络状态表示
- 在真实感环境中验证,防御性能不变且动作效率提升
- 单一模型可跨不同拓扑和规模网络部署,适合实际应用
深度强化学习在模拟网络防御中已取得进展,但现有智能体在面对不同拓扑或规模的网络时需重新训练,难以适应现实网络的动态变化。本文提出拓扑扩展强化学习智能体(TERLA),通过异构图神经网络层生成固定尺寸的网络状态嵌入,并结合语义明确、可解释的固定动作空间,实现无需重训练的泛化防御能力。将TERLA与标准PPO算法结合,使用Cyber Autonomy Gym for Experimentation(CAGE)挑战4环境进行测试,该环境具备真实网络特征,如真实入侵检测系统事件及多智能体协同防御不同规模网络段的能力。实验表明,所有TERLA智能体采用相同的网络无关架构,单个模型多次部署于不同拓扑和规模的网络段,均保持原有防御性能并提升动作效率,显著缩小了仿真到现实的差距。
原文摘要 · Abstract (English)
Recent advances in deep reinforcement learning for autonomous cyber defence have resulted in agents that can successfully defend simulated computer networks against cyber-attacks. However, many of these agents would need retraining to defend networks with differing topology or size, making them poorly suited to real-world networks where topology and size can vary over time. In this research we introduce a novel set of Topological Extensions for Reinforcement Learning Agents (TERLA) that provide generalisability for the defence of networks with differing topology and size, without the need for retraining. Our approach involves the use of heterogeneous graph neural network layers to produce a fixed-size latent embedding representing the observed network state. This representation learning stage is coupled with a reduced, fixed-size, semantically meaningful and interpretable action space. We apply TERLA to a standard deep reinforcement learning Proximal Policy Optimisation (PPO) agent model, and to reduce the sim-to-real gap, conduct our research using Cyber Autonomy Gym for Experimentation (CAGE) Challenge 4. This Cyber Operations Research Gym environment has many of the features of a real-world network, such as realistic Intrusion Detection System (IDS) events and multiple agents defending network segments of differing topology and size. TERLA agents retain the defensive performance of vanilla PPO agents whilst showing improved action efficiency. Generalisability has been demonstrated by showing that all TERLA agents have the same network-agnostic neural network architecture, and by deploying a single TERLA agent multiple times to defend network segments with differing topology and size, showing improved defensive performance and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。