arXiv:2607.18359cs.MAcs.LG2026-07中稿 · the 2026 IEEE Inte…

用去中心化强化学习提升关键基础设施的抗灾能力

Decentralized Multi-agent Reinforcement Learning for Resilient Critical Infrastructures

论文配图:Decentralized Multi-agent Reinforcement Learning for Resilient Critical Infrastructures
图 1 · 摘自论文原文
  • 将多智能体强化学习结构与抗灾需求对齐
  • 解决局部学习与全局目标一致性的信用分配难题
  • 适合关注基础设施安全与自主协同的研究者

关键基础设施日益分布化、相互依赖,并面临不断演变的中断威胁,韧性成为其运行与控制的核心要求。本文认为,去中心化多智能体强化学习(MARL)不应仅被视为集中训练、分散执行的替代方案,而应作为与韧性基础设施需求在结构上契合的范式。这一观点基于对去中心化MARL特性与基础设施需求(如大规模可扩展性、隐私保护、本地自治、容错能力,以及组件间互动驱动的自适应)的分析。然而,结构契合不足以支撑实际部署。本文指出信用分配与通信是其实用可行性的两大核心条件:前者决定局部学习是否与系统级目标对齐,后者决定在真实约束下能否学习并维持协调。据此,本文提出研究议程,聚焦于结构感知、因果感知和韧性感知的信用分配;支持协调与信用分配的通信机制;以及在部署约束下的安全、及时、可恢复的去中心化学习。总体而言,本文将去中心化MARL重构为韧性基础设施的有前景但有条件的基础。

原文摘要 · Abstract (English)

Critical infrastructures are increasingly distributed, interdependent, and exposed to evolving disruptions, making resilience a central requirement for their operation and control. This paper argues that decentralized multi-agent reinforcement learning (MARL) should be understood not merely as a distributed alternative to centralized training with decentralized execution but as a paradigm structurally aligned with the requirements of resilient critical infrastructures. This perspective is grounded in an analysis of the properties of decentralized MARL and the requirements of critical infrastructures, including scalability to large numbers of agents, support for privacy and local autonomy, robustness to failures, and interaction-driven adaptation among interdependent components. However, structural alignment alone is insufficient for practical deployment. This paper identifies credit assignment and communication as two central conditions for its practical feasibility. Credit assignment determines whether local learning remains aligned with system-level objectives, while communication determines whether coordination can be learned and maintained under realistic operational constraints. Building on these challenges, this paper proposes a research agenda focused on structure-aware, causality-aware, and resilience-aware credit assignment; communication for both coordination and credit assignment; and safe, timely, and recoverable decentralized learning under deployment constraints. Overall, this paper reframes decentralized MARL as a promising but conditional foundation for resilient critical infrastructures.

强化学习基础设施去中心化韧性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。