arXiv:2511.13103cs.LGcs.MA2025-11被引 1

用Transformer解决网络系统中远距离交互与跨拓扑泛化难题

Transformer-Based Multi-Agent Reinforcement Learning for Networked Systems with Long-Range Interactions

  • 基于图Transformer的集中式批评者建模长程依赖,共享演员适应多种网络结构
  • 在疫情控制任务中相较基线提升37%的覆盖率,且在未见拓扑上保持性能
  • 适用于大规模网络控制场景,尤其适合存在远距离传播的系统

多智能体强化学习(MARL)在大规模网络控制中展现出潜力,但现有方法存在两大局限:其一,依赖远端节点间交互呈指数衰减的假设,当该假设不成立时(如疫情传播或电力故障连锁反应),无法有效处理长程交互;其二,缺乏对不同拓扑结构和规模网络的泛化能力。本文通过均值场稳定性分析与实证研究,揭示了长程交互的形成条件,并提出STACCA(共享Transformer演员-评论家与反事实优势估计)框架。该框架采用中心化图Transformer评论家建模长程依赖并提供系统级反馈,共享图Transformer演员学习可泛化的策略以适应多样网络拓扑。为改善信用分配,引入兼容状态值评论家的新型反事实优势估计器。在疫情遏制与谣言传播控制任务中,STACCA显著优于基线方法,在未见网络结构上仍保持良好性能,验证了基于Transformer的MARL在大规模网络系统中实现可泛化控制的可行性。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations. First, they typically rely on an exponential decay property of agent interactions on far-away nodes, which can be exploited to develop more efficient and tractable MARL algorithms. When this exponential decay property does not hold, these algorithms do not account for long-range interactions such as epidemic outbreaks or cascading power failures. Second, existing approaches lack network generalizability, or the ability to generalize to networks of different topological structure and scale than those seen during training. In this work, we first present a mean-field stability analysis and empirical study investigating the conditions for long-range network interactions. These results motivate our primary contribution: STACCA (Shared Transformer Actor-Critic with Counterfactual Advantage), a transformer-based MARL framework that addresses both long-range interactions and network generalizability. STACCA employs a centralized Graph Transformer Critic to model long-range dependencies and provide system-level feedback, while its shared Graph Transformer Actor learns a generalizable policy capable of adapting across diverse network topologies. To improve credit assignment during training, STACCA integrates a novel counterfactual advantage estimator that is compatible with state-value critic estimates. We evaluate STACCA on epidemic containment and rumor-spreading network control tasks, demonstrating improved performance and network generalizability. These results highlight the potential of transformer-based MARL architectures to achieve generalizable control in large-scale networked systems.

多智能体强化学习图神经网络网络控制Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。