稀疏奖励比密集奖励更利于训练出安全可靠的自适应网络防御智能体。
Beyond Rewards in Reinforcement Learning for Cyber Defence
- 用稀疏奖励替代密集奖励,避免诱导次优策略。
- 稀疏奖励下防御策略风险更低,且减少高成本防御动作使用。
- 适用于追求高可靠性和低风险的网络安全系统设计者。
近年来,基于深度强化学习的自主网络防御智能体受到广泛关注。这些智能体通常在网络安全模拟环境(cyber gym)中使用密集、高度工程化的奖励函数进行训练,该函数结合了多种对(不)期望状态和高成本动作的惩罚与激励。虽然密集奖励有助于缓解复杂环境中的探索难题,但可能使智能体偏向次优甚至更危险的策略,这在复杂的网络环境中尤为关键。本文通过多种稀疏与密集奖励函数,在两个成熟的网络安全模拟环境、多种网络规模以及基于策略梯度和值函数的强化学习算法下,系统评估了奖励结构对学习过程和策略行为的影响。评估依托一种新型真实基准测试方法,可直接比较不同奖励函数的表现,揭示了奖励、动作空间与次优策略风险之间的复杂关系。结果表明,只要目标对齐且能频繁触发,稀疏奖励不仅能提升训练稳定性,还能生成风险更低、更符合防御目标的智能体,且无需显式数值惩罚即可减少高成本防御动作的使用。
原文摘要 · Abstract (English)
Recent years have seen an explosion of interest in autonomous cyber defence agents trained to defend computer networks using deep reinforcement learning. These agents are typically trained in cyber gym environments using dense, highly engineered reward functions which combine many penalties and incentives for a range of (un)desirable states and costly actions. Dense rewards help alleviate the challenge of exploring complex environments but risk biasing agents towards suboptimal and potentially riskier solutions, a critical issue in complex cyber environments. We thoroughly evaluate the impact of reward function structure on learning and policy behavioural characteristics using a variety of sparse and dense reward functions, two well-established cyber gyms, a range of network sizes, and both policy gradient and value-based RL algorithms. Our evaluation is enabled by a novel ground truth evaluation approach which allows directly comparing between different reward functions, illuminating the nuanced inter-relationships between rewards, action space and the risks of suboptimal policies in cyber environments. Our results show that sparse rewards, provided they are goal aligned and can be encountered frequently, uniquely offer both enhanced training reliability and more effective cyber defence agents with lower-risk policies. Surprisingly, sparse rewards can also yield policies that are better aligned with cyber defender goals and make sparing use of costly defensive actions without explicit reward-based numerical penalties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。