用课程学习和动作掩码提升工业网络安全强化学习的数据效率与性能。
Applying Action Masking and Curriculum Learning Techniques to Improve Data Efficiency and Overall Performance in Operational Technology Cyber Security using Reinforcement Learning
- 引入课程学习与动作掩码,分阶段训练智能体应对复杂攻击场景。
- 组合使用使平均奖励达0.137,在百万步内完成训练,远超基准方法。
- 适合研究工业网络安全、强化学习高效训练的学者与工程师。
先前工作开发了IPMSRL环境(集成平台管理系统强化学习环境),用于在模拟海事船舶部分IPMS系统受网络攻击时训练防御型强化学习代理。本文扩展了该环境以增强现实性,加入误报警动态与告警延迟。在最困难环境下,应用课程学习使每回合平均奖励从基线-2.791提升至-0.569;应用动作掩码则提升至-0.743。重要的是,该性能在不足100万时间步内达成,远优于原生PPO在250万时间步后取得的较低表现。最优方法为课程学习与动作掩码联合使用,实现平均奖励0.137。本文还引入一个硬编码防御代理,代表网络安全最佳实践,其平均奖励为-1.895。结果表明,课程学习与动作掩码的独立或协同应用,是应对工业网络安全威胁修复中复杂现实动态的有效手段。
原文摘要 · Abstract (English)
In previous work, the IPMSRL environment (Integrated Platform Management System Reinforcement Learning environment) was developed with the aim of training defensive RL agents in a simulator representing a subset of an IPMS on a maritime vessel under a cyber-attack. This paper extends the use of IPMSRL to enhance realism including the additional dynamics of false positive alerts and alert delay. Applying curriculum learning, in the most difficult environment tested, resulted in an episode reward mean increasing from a baseline result of -2.791 to -0.569. Applying action masking, in the most difficult environment tested, resulted in an episode reward mean increasing from a baseline result of -2.791 to -0.743. Importantly, this level of performance was reached in less than 1 million timesteps, which was far more data efficient than vanilla PPO which reached a lower level of performance after 2.5 million timesteps. The training method which resulted in the highest level of performance observed in this paper was a combination of the application of curriculum learning and action masking, with a mean episode reward of 0.137. This paper also introduces a basic hardcoded defensive agent encoding a representation of cyber security best practice, which provides context to the episode reward mean figures reached by the RL agents. The hardcoded agent managed an episode reward mean of -1.895. This paper therefore shows that applications of curriculum learning and action masking, both independently and in tandem, present a way to overcome the complex real-world dynamics that are present in operational technology cyber security threat remediation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。