arXiv:2603.16969cs.CRcs.AI2026-03中稿 · publication in IEE…被引 2

用深度强化学习实现分阶段自适应防御,提升对高级持续性威胁的拦截能力。

DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns

  • 基于图神经网络与LSTM推断攻击阶段,融合主机和网络数据
  • 在真实企业环境中实现84.7%的缓解成功率和0.887的F1分数
  • 适合需要自动化、分阶段响应的网络安全团队使用

本文提出DeepStage,一种针对高级持续性威胁(APT)的深度强化学习(DRL)框架,实现自适应且阶段感知的防御。将企业环境建模为部分可观测马尔可夫决策过程(POMDP),通过融合主机溯源与网络遥测数据生成统一的溯源图。基于前期工作StageFinder,DeepStage采用图神经网络编码器和基于LSTM的阶段估计器,推断与MITRE ATT&CK框架对齐的攻击阶段概率。由此获得的阶段信念与图嵌入共同指导分层近端策略优化(PPO)智能体,在监控、访问控制、隔离和修复等层面选择防御动作。在基于CALDERA驱动的APT剧本的真实企业测试平台上实验表明,DeepStage平均F1得分为0.887,缓解成功率达84.7%,相比风险感知型DRL基线提升21.8%(F1)和16.2%(缓解成功率)。结果证明该方法能有效实现阶段感知且成本高效的自主网络防御。

原文摘要 · Abstract (English)

This paper presents DeepStage, a deep reinforcement learning (DRL) framework for adaptive and stage-aware defense against Advanced Persistent Threats (APTs). The enterprise environment is formulated as a partially observable Markov decision process (POMDP), in which host provenance and network telemetry are fused into unified provenance graphs. Building on our prior work (StageFinder), DeepStage employs a graph neural network encoder and an LSTM-based stage estimator to infer probabilistic attacker stages aligned with the MITRE ATT&CK framework. The resulting stage beliefs, together with graph embeddings, are used to guide a hierarchical Proximal Policy Optimization (PPO) agent that selects defense actions across monitoring, access control, containment, and remediation. Experiments in a realistic enterprise testbed with CALDERA-driven APT playbooks show that DeepStage achieves an average F1-score of 0.887 and a mitigation success rate of 84.7%, outperforming a risk-aware DRL baseline by 21.8% in F1-score and 16.2% in mitigation success. The results demonstrate effective stage-aware and cost-efficient autonomous cyber defense.

强化学习APT防御图神经网络自主安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。