arXiv:2510.07363cs.AI2025-10中稿 · version will be pu…被引 6

用大模型+多智能体强化学习实现工业系统自主防御,兼顾安全与生产稳定。

L2M-AID: Autonomous Cyber-Physical Defense by Fusing Semantic Reasoning of Large Language Models with Multi-Agent Reinforcement Learning (Preprint)

  • 大模型解析日志生成语义状态,让智能体理解攻击意图而非仅识别模式
  • 在SWaT数据集上检测率达97.2%,误报率降低80%,响应速度提升4倍
  • 适合关注工控安全、智能防御系统的研究人员和工程师

工业物联网(IIoT)的普及使关键网络物理系统面临复杂多阶段攻击,传统防御因缺乏上下文感知而失效。本文提出L2M-AID框架,融合大语言模型(LLM)与多智能体强化学习(MARL)实现自主工业防御。核心创新在于利用LLM作为语义桥梁,将海量非结构化遥测数据转化为富含上下文的状态表示,使智能体能推理攻击者意图,而非仅匹配异常模式。该语义状态驱动基于MAPPO的多智能体强化学习,其奖励函数同时平衡安全目标(威胁消除)与运行需求,明确惩罚影响物理过程稳定性的动作。我们在标准SWaT数据集及基于MITRE ATT&CK for ICS构建的新合成数据集上进行验证。结果表明,L2M-AID显著优于传统IDS、深度学习异常检测器及单智能体强化学习基线,在检测率(97.2%)、误报率(降低超80%)和响应时间(提升4倍)等指标上全面领先,并有效维持了物理过程稳定性,为关键基础设施安全提供了可靠新范式。

原文摘要 · Abstract (English)

The increasing integration of Industrial IoT (IIoT) exposes critical cyber-physical systems to sophisticated, multi-stage attacks that elude traditional defenses lacking contextual awareness. This paper introduces L2M-AID, a novel framework for Autonomous Industrial Defense using LLM-empowered, Multi-agent reinforcement learning. L2M-AID orchestrates a team of collaborative agents, each driven by a Large Language Model (LLM), to achieve adaptive and resilient security. The core innovation lies in the deep fusion of two AI paradigms: we leverage an LLM as a semantic bridge to translate vast, unstructured telemetry into a rich, contextual state representation, enabling agents to reason about adversary intent rather than merely matching patterns. This semantically-aware state empowers a Multi-Agent Reinforcement Learning (MARL) algorithm, MAPPO, to learn complex cooperative strategies. The MARL reward function is uniquely engineered to balance security objectives (threat neutralization) with operational imperatives, explicitly penalizing actions that disrupt physical process stability. To validate our approach, we conduct extensive experiments on the benchmark SWaT dataset and a novel synthetic dataset generated based on the MITRE ATT&CK for ICS framework. Results demonstrate that L2M-AID significantly outperforms traditional IDS, deep learning anomaly detectors, and single-agent RL baselines across key metrics, achieving a 97.2% detection rate while reducing false positives by over 80% and improving response times by a factor of four. Crucially, it demonstrates superior performance in maintaining physical process stability, presenting a robust new paradigm for securing critical national infrastructure.

工控安全大模型多智能体自主防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。