稀疏奖励让网络防御智能体更有效且训练更稳定
Less is more? Rewards in RL for Cyber Defence
- 用稀疏奖励替代密集奖励,聚焦网络未被攻破的状态
- 在2到50个节点的网络中,稀疏奖励提升防御效果和训练稳定性
- 适合研究强化学习在网络安全中的应用或想避免奖励偏差的研究者
近年来,基于深度强化学习的自主网络安全防御代理受到广泛关注。这些代理通常在至少32个已构建的网络模拟环境(即网络模拟器)中训练。大多数甚至所有网络模拟器都采用密集的“结构化”奖励函数,结合多种惩罚与激励以应对各类(不)理想状态及高成本操作。尽管密集奖励有助于缓解复杂环境下的探索难题,使代理在较少环境步数下产生看似有效的策略,但它们也可能导致解空间偏差,趋向次优解。尤其在复杂网络环境中,策略缺陷可能直到被对手利用才被发现。本文旨在评估稀疏奖励是否能训练出更有效的网络安全防御代理。为此,我们提出一个超越传统强化学习范式的基准评估分数,并改造成熟网络模拟器以支持该方法。我们设计并评估了两种稀疏奖励机制,与典型密集奖励进行对比。实验覆盖2至50个节点的网络规模,涵盖反应式与主动防御行为。结果表明,稀疏奖励(尤其是对未被攻破网络状态的正向激励)能训练出更有效的防御代理;同时,稀疏奖励提供更稳定的训练过程,且有效性与稳定性在多种网络环境下均表现稳健。
原文摘要 · Abstract (English)
The last few years have seen an explosion of interest in autonomous cyber defence agents based on deep reinforcement learning. Such agents are typically trained in a cyber gym environment, also known as a cyber simulator, at least 32 of which have already been built. Most, if not all cyber gyms provide dense "scaffolded" reward functions which combine many penalties or incentives for a range of (un)desirable states and costly actions. Whilst dense rewards help alleviate the challenge of exploring complex environments, yielding seemingly effective strategies from relatively few environment steps; they are also known to bias the solutions an agent can find, potentially towards suboptimal solutions. This is especially a problem in complex cyber environments where policy weaknesses may not be noticed until exploited by an adversary. In this work we set out to evaluate whether sparse reward functions might enable training more effective cyber defence agents. Towards this goal we first break down several evaluation limitations in existing work by proposing a ground truth evaluation score that goes beyond the standard RL paradigm used to train and evaluate agents. By adapting a well-established cyber gym to accommodate our methodology and ground truth score, we propose and evaluate two sparse reward mechanisms and compare them with a typical dense reward. Our evaluation considers a range of network sizes, from 2 to 50 nodes, and both reactive and proactive defensive actions. Our results show that sparse rewards, particularly positive reinforcement for an uncompromised network state, enable the training of more effective cyber defence agents. Furthermore, we show that sparse rewards provide more stable training than dense rewards, and that both effectiveness and training stability are robust to a variety of cyber environment considerations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。