通过模仿学习预测红方攻击行为,提升防御智能体的应对能力
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents

- 用模仿学习从网络观测中推断不可见的红方策略
- 在多种模拟场景下实现高精度攻击行为预测
- 适合需要动态对抗的网络安全系统开发者
随着复杂网络攻击日益普遍,现代网络需要通过强化学习训练的智能自主防御代理。这些代理采用神经符号方法,如带有学习增强组件(LECs)的行为树,以学习、推理、适应并执行安全规则,同时维持关键操作。然而,这些自治网络是部分可观测系统,即无法直接观察攻击者(红方)的行为,导致防御者难以预测红方动作、学习其策略或评估入侵程度。为此,我们提出一种针对具有离散状态和离散动作的部分可观测强化学习代理的策略学习技术,采用模仿学习方法。该方法在自主网络环境中应用,能够根据网络观测和防御者动作预测红方行为。与神经符号防御代理集成后,该方法能有效应对不同红方策略,并在多样化的模拟场景中实现高预测准确率。
原文摘要 · Abstract (English)
With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents trained via Reinforcement Learning (RL). These agents employ neurosymbolic approaches such as behavior trees with learning-enabled components (LECs) to learn, reason, adapt, and implement security rules while maintaining critical operations. However, these autonomous networks are partially observable systems, i.e., the cyber-attacker's (red agent's) actions are not observable, making it difficult for the defender to predict red actions, learn red policies, or assess the attacker's intrusion levels. To address this, we propose a Policy Learning Technique using imitation learning to learn policies for partially observable RL agents with discrete states and discrete actions. We apply this technique in an autonomous cyber environment to predict red agent's actions from network observations and defender actions. Integrated with a neurosymbolic cyber-defense agent, our method effectively handles different red policies and achieves high prediction accuracy across diverse simulated scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。