用强化学习优化网络安全防御策略,效率比现有方法更高。
Learning Optimal Defender Strategies for CAGE-2 using a POMDP Model
- 基于部分可观测马尔可夫决策过程构建防御模型
- 新方法BF-PPO在训练时间和防御效果上均优于顶尖方案CARDIFF
- 适合研究智能网络安全与强化学习应用的学者参考
CAGE-2是评估网络防御策略的公认基准,模拟防御者保护信息系统免受各类攻击的场景。本文采用部分可观测马尔可夫决策过程(POMDP)框架,为CAGE-2构建形式化模型,并定义了最优防御策略。提出一种名为BF-PPO的方法,基于PPO算法并结合粒子滤波以缓解大规模状态空间带来的计算复杂性。在CAGE-2 CybORG环境中评估该方法,结果表明其在学习到的防御策略性能和训练时间方面均优于当前排行榜领先的CARDIFF方法。
原文摘要 · Abstract (English)
CAGE-2 is an accepted benchmark for learning and evaluating defender strategies against cyberattacks. It reflects a scenario where a defender agent protects an IT infrastructure against various attacks. Many defender methods for CAGE-2 have been proposed in the literature. In this paper, we construct a formal model for CAGE-2 using the framework of Partially Observable Markov Decision Process (POMDP). Based on this model, we define an optimal defender strategy for CAGE-2 and introduce a method to efficiently learn this strategy. Our method, called BF-PPO, is based on PPO, and it uses particle filter to mitigate the computational complexity due to the large state space of the CAGE-2 model. We evaluate our method in the CAGE-2 CybORG environment and compare its performance with that of CARDIFF, the highest ranked method on the CAGE-2 leaderboard. We find that our method outperforms CARDIFF regarding the learned defender strategy and the required training time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。