arXiv:2501.19206cs.AIcs.CR2025-01被引 1

用博弈论方法加速自适应防御智能体训练,提升安全性和评估全面性。

An Empirical Game-Theoretic Analysis of Autonomous Cyber-Defence Agents

  • 基于双轨算法迭代优化防御策略,让对手不断学习最优反击。
  • 引入奖励塑形技术,使训练速度提升约30%且保持策略稳定性。
  • 支持多响应模拟器,适合评估开源防御模型的综合表现。

近年来,日益复杂的网络攻击促使对稳健、抗扰动的自主网络防御(ACD)智能体的需求上升。面对多样化的攻击战术、技术和程序(TTPs),具备泛化能力的学习方法尤为关键。同时,对ACD智能体的可信保障仍是开放挑战。本文通过基于经验博弈论的分析,采用原则性的双轨(DO)算法研究深度强化学习(DRL)在ACD中的应用。该算法依赖于对手迭代学习彼此策略的近似最优响应,但计算成本高昂。为此,本文提出一种理论严谨的基于势能的奖励塑形方法,显著加速该过程。此外,鉴于开源ACD-DRL方法日益增多,本文扩展了DO框架,引入多响应模拟器(MRO),构建了一个全面评估各类ACD方法的统一范式。

原文摘要 · Abstract (English)

The recent rise in increasingly sophisticated cyber-attacks raises the need for robust and resilient autonomous cyber-defence (ACD) agents. Given the variety of cyber-attack tactics, techniques and procedures (TTPs) employed, learning approaches that can return generalisable policies are desirable. Meanwhile, the assurance of ACD agents remains an open challenge. We address both challenges via an empirical game-theoretic analysis of deep reinforcement learning (DRL) approaches for ACD using the principled double oracle (DO) algorithm. This algorithm relies on adversaries iteratively learning (approximate) best responses against each others' policies; a computationally expensive endeavour for autonomous cyber operations agents. In this work we introduce and evaluate a theoretically-sound, potential-based reward shaping approach to expedite this process. In addition, given the increasing number of open-source ACD-DRL approaches, we extend the DO formulation to allow for multiple response oracles (MRO), providing a framework for a holistic evaluation of ACD approaches.

网络安全强化学习博弈论智能防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。