让自动防御系统能解释决策,提升对复杂攻击的应对可信度。
DeepXplain: XAI-Guided Autonomous Defense Against Multi-Stage APT Campaigns
- 用可解释AI直接优化防御策略,结合时间阶段与溯源图学习。
- 防御准确率提升至89.6%,解释可信度达0.86,更简洁可靠。
- 适合需要高可信自动化防御的金融、能源等关键行业。
高级持续性威胁(APTs)是隐蔽且多阶段的攻击,需具备自适应与及时响应能力。尽管深度强化学习(DRL)可实现自主防御,但其决策过程常不透明,难以在实际环境中获得信任。本文提出DeepXplain,一种面向阶段感知的APT防御可解释深度强化学习框架。基于先前的DeepStage模型,DeepXplain融合溯源图学习、时间阶段估计及统一的XAI流程,提供结构、时间与策略层面的解释。不同于事后解释方法,解释信号通过证据对齐与置信度感知奖励设计,直接嵌入策略优化过程。据我们所知,DeepXplain是首个将解释信号融入DRL以应对APT防御的框架。在真实企业测试环境中,其阶段加权F1得分从0.887提升至0.915,成功率由84.7%增至89.6%,解释置信度达0.86,保真度为0.79,解释紧凑性为0.31。结果表明,该框架显著提升了自主防御的有效性与可信度。
原文摘要 · Abstract (English)
Advanced Persistent Threats (APTs) are stealthy, multi-stage attacks that require adaptive and timely defense. While deep reinforcement learning (DRL) enables autonomous cyber defense, its decisions are often opaque and difficult to trust in operational environments. This paper presents DeepXplain, an explainable DRL framework for stage-aware APT defense. Building on our prior DeepStage model, DeepXplain integrates provenance-based graph learning, temporal stage estimation, and a unified XAI pipeline that provides structural, temporal, and policy-level explanations. Unlike post-hoc methods, explanation signals are incorporated directly into policy optimization through evidence alignment and confidence-aware reward shaping. To the best of our knowledge, DeepXplain is the first framework to integrate explanation signals into reinforcement learning for APT defense. Experiments in a realistic enterprise testbed show improvements in stage-weighted F1-score (0.887 to 0.915) and success rate (84.7% to 89.6%), along with higher explanation confidence (0.86), improved fidelity (0.79), and more compact explanations (0.31). These results demonstrate enhanced effectiveness and trustworthiness of autonomous cyber defense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。