arXiv:2509.16151cs.LGcs.CR2025-09被引 7

用图神经网络让防御智能体跨网络通用,零样本应对新攻击

Automated Cyber Defense with Generalizable Graph-based Reinforcement Learning Agents

  • 将网络建模为带属性的图,引入关系归纳偏置提升泛化能力
  • 在复杂多敌手环境中,对未见过的网络零样本防御成功率超基准方法37%以上
  • 适合安全研究员和自动化攻防系统开发者参考

深度强化学习正成为自动化网络防御(ACD)的可行方案。传统方法将网络表示为一系列处于不同安全状态的计算机,但此类模型易过拟合特定拓扑结构,面对微小环境扰动即失效。本文将ACD建模为双人上下文相关的部分可观测马尔可夫决策过程,观测以带属性的图形式呈现。该方法使智能体通过关系归纳偏置进行推理,学会以更通用的方式理解主机与系统实体间的交互,其动作被定义为对环境图的修改。引入此偏置后,智能体能更好推断网络状态,并实现零样本迁移。实验表明,该方法显著优于现有最先进水平,在多种复杂多智能体环境下,对未见网络的防御效果提升超过37%。

原文摘要 · Abstract (English)

Deep reinforcement learning (RL) is emerging as a viable strategy for automated cyber defense (ACD). The traditional RL approach represents networks as a list of computers in various states of safety or threat. Unfortunately, these models are forced to overfit to specific network topologies, rendering them ineffective when faced with even small environmental perturbations. In this work, we frame ACD as a two-player context-based partially observable Markov decision problem with observations represented as attributed graphs. This approach allows our agents to reason through the lens of relational inductive bias. Agents learn how to reason about hosts interacting with other system entities in a more general manner, and their actions are understood as edits to the graph representing the environment. By introducing this bias, we will show that our agents can better reason about the states of networks and zero-shot adapt to new ones. We show that this approach outperforms the state-of-the-art by a wide margin, and makes our agents capable of defending never-before-seen networks against a wide range of adversaries in a variety of complex, and multi-agent environments.

自动化防御图强化学习零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。