arXiv:2601.07122cs.CRcs.AI2026-01

用大模型+多智能体强化学习提升云网络抗攻击能力,无需重训即可适应变化。

Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework

  • 大模型负责全局规划与人类意图理解,底层智能体执行具体防御动作
  • 在不重训练情况下,网络可用性提升68.5%,响应速度加快34.7%
  • 支持人机协同,可解释性强,适合动态云安全场景

虚拟化与资源池化虽赋予云网络弹性扩展能力,却也扩大了攻击面,挑战网络安全韧性。现有基于强化学习的防御方法虽能优化资源配置与隔离策略,但在面对网络结构、节点规模、攻击策略及强度动态变化时缺乏鲁棒性,且缺少人机协同支持,导致可解释性与灵活性不足。为此,本文提出CyberOps-Bots——一种由大语言模型赋能的分层多智能体强化学习框架。该框架受MITRE ATT&CK战术-技术模型启发,采用双层架构:上层为具四模块(ReAct规划、IPDRR感知、长短时记忆、动作/工具融合)的LLM智能体,实现全局态势感知、人类意图识别与战术规划;下层为通过异构分离预训练的多个强化学习智能体,在局部区域执行原子级防御操作。二者协同兼顾大模型的适应性与可解释性,同时保障强化学习的可靠执行。在真实云数据集上的实验表明,相比当前最优算法,该框架在不重训练的前提下,网络可用性提升68.5%,性能启动加速34.7%。据我们所知,这是首个具备人机协同支持的鲁棒型大模型-强化学习云防御框架。

原文摘要 · Abstract (English)

While virtualization and resource pooling empower cloud networks with structural flexibility and elastic scalability, they inevitably expand the attack surface and challenge cyber resilience. Reinforcement Learning (RL)-based defense strategies have been developed to optimize resource deployment and isolation policies under adversarial conditions, aiming to enhance system resilience by maintaining and restoring network availability. However, existing approaches lack robustness as they require retraining to adapt to dynamic changes in network structure, node scale, attack strategies, and attack intensity. Furthermore, the lack of Human-in-the-Loop (HITL) support limits interpretability and flexibility. To address these limitations, we propose CyberOps-Bots, a hierarchical multi-agent reinforcement learning framework empowered by Large Language Models (LLMs). Inspired by MITRE ATT&CK's Tactics-Techniques model, CyberOps-Bots features a two-layer architecture: (1) An upper-level LLM agent with four modules--ReAct planning, IPDRR-based perception, long-short term memory, and action/tool integration--performs global awareness, human intent recognition, and tactical planning; (2) Lower-level RL agents, developed via heterogeneous separated pre-training, execute atomic defense actions within localized network regions. This synergy preserves LLM adaptability and interpretability while ensuring reliable RL execution. Experiments on real cloud datasets show that, compared to state-of-the-art algorithms, CyberOps-Bots maintains network availability 68.5% higher and achieves a 34.7% jumpstart performance gain when shifting the scenarios without retraining. To our knowledge, this is the first study to establish a robust LLM-RL framework with HITL support for cloud defense.

云安全多智能体大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。