arXiv:2508.19488cs.LGcs.AI2025-08中稿 · GameSec 2025被引 1

用多智能体强化学习打造更智能的网络防御系统。

PoolFlip: A Multi-Agent Reinforcement Learning Security Environment for Cyber Defense

  • 构建扩展版FlipIt游戏环境,支持攻防双方智能体学习。
  • 新方法使防御方在未知攻击下效果提升两倍。
  • 适合研究网络安全对抗与智能防御的学者。

网络防御需要在隐蔽、欺骗性和持续演化的攻击策略下自动化决策。经典FlipIt模型通过攻防双方争夺共享资源来模拟这种对抗,但现有框架依赖少量启发式规则或专用学习技术,易导致脆弱性且难以适应新攻击。为此,我们提出PoolFlip,一个扩展的多智能体强化学习环境,可高效训练攻防双方智能体。同时,我们设计了基于种群的强化学习方法Flip-PSRO,利用群体训练使防御智能体能泛化应对多种未知、可能自适应的对手。实验表明,采用Flip-PSRO的防御者在未训练过的启发式攻击下表现提升2倍。此外,新设计的基于所有权的效用函数确保防御者在优化性能的同时保持高控制水平。

原文摘要 · Abstract (English)

Cyber defense requires automating defensive decision-making under stealthy, deceptive, and continuously evolving adversarial strategies. The FlipIt game provides a foundational framework for modeling interactions between a defender and an advanced adversary that compromises a system without being immediately detected. In FlipIt, the attacker and defender compete to control a shared resource by performing a Flip action and paying a cost. However, the existing FlipIt frameworks rely on a small number of heuristics or specialized learning techniques, which can lead to brittleness and the inability to adapt to new attacks. To address these limitations, we introduce PoolFlip, a multi-agent gym environment that extends the FlipIt game to allow efficient learning for attackers and defenders. Furthermore, we propose Flip-PSRO, a multi-agent reinforcement learning (MARL) approach that leverages population-based training to train defender agents equipped to generalize against a range of unknown, potentially adaptive opponents. Our empirical results suggest that Flip-PSRO defenders are $2\times$ more effective than baselines to generalize to a heuristic attack not exposed in training. In addition, our newly designed ownership-based utility functions ensure that Flip-PSRO defenders maintain a high level of control while optimizing performance.

网络安全多智能体强化学习防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。