arXiv:2505.11708cs.CRcs.LG2025-05被引 5

为强化学习攻击代理设计多层可解释框架,揭示其攻防策略演化过程。

Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents

  • 将攻击建模为部分可观测马尔可夫决策过程,分析探索与利用动态及阶段行为变化。
  • 通过Q值时序演化和优先经验回放,识别关键学习转折点与行动偏好转变。
  • 适用于红队演练、策略调试等场景,支持防御方提前预判自主威胁。

强化学习(RL)代理被广泛用于模拟复杂网络攻击,但其决策过程高度黑箱,影响信任建立、故障排查与防御准备。在高风险网络安全场景中,可解释性对理解对抗策略的形成与演化至关重要。本文提出一个统一的多层可解释性框架,用于揭示基于强化学习的攻击代理在策略层(MDP级)与战术层(策略级)的推理机制。在MDP级,将攻击建模为部分可观测马尔可夫决策过程(POMDP),暴露探索-利用动态及阶段感知的行为转变;在策略级,分析Q值的时间演化,并利用优先经验回放(PER)揭示关键学习转折点与持续演变的动作偏好。在复杂度递增的CyberBattleSim环境中评估,该框架实现了大规模下对代理行为的可解释洞察。与以往主要为事后、领域特定或深度有限的可解释方法不同,本方法具有代理与环境无关性,支持红队模拟、策略调试、阶段感知威胁建模及前瞻防御规划。通过将黑箱学习转化为可操作的行为智能,助力防御者与开发者更有效地预见、分析与响应自主网络威胁。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) agents are increasingly used to simulate sophisticated cyberattacks, but their decision-making processes remain opaque, hindering trust, debugging, and defensive preparedness. In high-stakes cybersecurity contexts, explainability is essential for understanding how adversarial strategies are formed and evolve over time. In this paper, we propose a unified, multi-layer explainability framework for RL-based attacker agents that reveals both strategic (Markov Decision Process (MDP)-level) and tactical (policy-level) reasoning. At the MDP-level, we model cyberattacks as a Partially Observable Markov Decision Process (POMDP) to expose exploration-exploitation dynamics and phase-aware behavioural shifts. At the policy-level, we analyse the temporal evolution of Q-values and use Prioritised Experience Replay (PER) to surface critical learning transitions and evolving action preferences. Evaluated across CyberBattleSim environments of increasing complexity, our framework offers interpretable insights into agent behaviour at scale. Unlike previous explainable RL methods, which are {predominantly} post-hoc, domain-specific, or limited in depth, our approach is both agent- and environment-agnostic, {supporting use cases such as red-team simulation, RL policy debugging, phase-aware threat modelling and anticipatory defence planning.} By transforming black-box learning into actionable behavioural intelligence, our framework enables both defenders and developers to better anticipate, analyse, and respond to autonomous cyber threats.

强化学习可解释性网络安全红队演练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。