arXiv:2606.13832cs.MAcs.AI2026-06

提出安全契约图强化学习框架,让网络安全自适应响应更可靠

Safety-Contract Graph Multi-Agent Reinforcement Learning for Autonomous Network Security Response

论文配图:Safety-Contract Graph Multi-Agent Reinforcement Learning for Autonomous Network Security Response
图 1 · 摘自论文原文
  • 用图注意力网络分离状态编码与约束优化,实现资源预算控制
  • 在测试中将系统宕机成本从355.4降至15.5,违规率从100%降到0.3%
  • 适合需要高可靠性、低误报的自动化安全运维场景

自主网络安防响应系统可降低安全运营中心(SOC)响应延迟,但仅靠奖励驱动的多智能体强化学习(MARL)虽能提升安全收益,却难以部署。本文提出安全契约图MARL框架,并实现为ACD³-GAT(自适应约束反事实决策,带图注意力网络编码器),该架构将仿真观测与可复用的操作预算、约束优化、图状态编码及反事实动作筛选相分离。在CAGE Challenge 4中评估,各智能体受限于平均恢复时间(MTTR)、误报率及防火墙变更扰动预算。所有无约束方法在100%测试回合中均超出SOC停机预算(预算为50),平均停机代理成本达311–430。相比之下,受约束的MAPPO-GAT(C-MAPPO-GAT)通过拉格朗日成本控制与预算感知筛选,将违规率降至0.3%,平均停机成本由355.4降至15.5。ACD³-GAT进一步引入预算上下文、CVaR尾部风险估计、对手信念状态与图反事实风险传播(G-CRP),使平均停机成本降至48.2,违规率为13.8%,位于安全契约前沿而非最保守点。拓扑种子与耦合自适应红队压力测试验证了该优势,且安全约束策略的最坏适应性退化低于奖励驱动的MAPPO-GAT。

原文摘要 · Abstract (English)

Autonomous network-security response systems promise to reduce Security Operations Centre (SOC) reaction latency, but reward-only multi-agent reinforcement learning (MARL) can improve security reward while remaining non-deployable. We present a safety-contract graph MARL framework and instantiate it as ACD$^3$-GAT (Adaptive Constrained Counterfactual Decisioning with a Graph Attention Network encoder), an architecture that separates simulator observations from reusable operational budgets, constrained optimization, graph state encoding, and counterfactual action screening. We evaluate the method in CAGE Challenge 4, where agents operate under budgets for Mean Time to Recover (MTTR), false-positive response, and firewall change-management disruption. Across the benchmark, every unconstrained method violates the SOC downtime budget in 100% of evaluated episodes, with mean downtime proxy costs of 311-430 against a budget of 50. This complements prior CAGE Challenge 4 findings by showing that reward-only learning lacks operational discipline. Constrained MAPPO-GAT (C-MAPPO-GAT) isolates Lagrangian operational-cost control and budget-aware screening, while ACD$^3$-GAT adds budget context, CVaR tail-risk estimation, opponent-belief state, and Graph Counterfactual Risk Propagation (G-CRP). The replicated comparison includes three 200-episode seeds for IPPO, MAPPO-GAT, C-MAPPO-GAT, and ACD$^3$-GAT. C-MAPPO-GAT reduces downtime violation from 100% to 0.3% and mean downtime cost from 355.4 to 15.5 relative to MAPPO-GAT. ACD$^3$-GAT reduces mean downtime cost to 48.2 with a 13.8% violation rate, placing it on the safety-contract frontier rather than at the most conservative compliance point. Topology-seed and coupled adaptive Red-process stress tests preserve this contrast and show lower worst adaptive degradation for safety-constrained policies than reward-only MAPPO-GAT.

网络安全强化学习多智能体约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。