arXiv:2503.05970cs.LGeess.SP2025-03

提出新型分布式多智能体强化学习算法,显著降低无线网络优化复杂度

Partially Decentralized Multi-Agent Q-Learning via Digital Cousins for Wireless Networks

  • 引入协同与非协同状态机制,分阶段优化本地与全局策略
  • 相比现有方法,平均策略误差降60%,收敛速度提升40%
  • 适合大规模分布式无线网络,兼顾性能与通信开销

Q-learning 广泛用于优化无线网络,但面临状态空间过大的挑战。近期提出的多环境混合 Q-learning(MEMQ)通过在多个结构相关但各不相同的合成环境(即数字孪生)中并行使用多个 Q-learning 算法来缓解该问题。本文提出一种新型多智能体 MEMQ(M-MEMQ),用于具有多个联网发射机(TXs)和基站(BSs)的协作式去中心化无线网络。TXs 无法获取全局信息(联合状态与动作)。本文引入协同与非协同状态的新概念:在非协同状态下,TXs 独立行动以最小化自身代价并更新本地 Q 函数;在协同状态下,TXs 使用贝叶斯方法估计联合状态并更新联合 Q 函数。信息共享成本随 TX 数线性增长,与联合状态-动作空间大小无关。本文提供了若干理论保证,包括确定性和概率收敛性、估计误差方差上界,以及误检联合状态的概率。数值仿真表明,M-MEMQ 在平均策略误差(APE)上优于多种去中心化与集中训练/去中心化执行(CTDE)的多智能体强化学习算法,分别实现 60% 降低、40% 更快收敛、45% 更低运行时复杂度和 40% 更少样本复杂度。此外,其性能接近集中式方法,但复杂度显著更低。仿真验证了理论分析。

原文摘要 · Abstract (English)

Q-learning is a widely used reinforcement learning (RL) algorithm for optimizing wireless networks, but faces challenges with large state-spaces. Recently proposed multi-environment mixed Q-learning (MEMQ) algorithm addresses these challenges by employing multiple Q-learning algorithms across multiple synthetically generated, distinct but structurally related environments, so-called digital cousins. In this paper, we propose a novel multi-agent MEMQ (M-MEMQ) for cooperative decentralized wireless networks with multiple networked transmitters (TXs) and base stations (BSs). TXs do not have access to global information (joint state and actions). The new concept of coordinated and uncoordinated states is introduced. In uncoordinated states, TXs act independently to minimize their individual costs and update local Q-functions. In coordinated states, TXs use a Bayesian approach to estimate the joint state and update the joint Q-functions. The cost of information-sharing scales linearly with the number of TXs and is independent of the joint state-action space size. Several theoretical guarantees, including deterministic and probabilistic convergence, bounds on estimation error variance, and the probability of misdetecting the joint states, are given. Numerical simulations show that M-MEMQ outperforms several decentralized and centralized training with decentralized execution (CTDE) multi-agent RL algorithms by achieving 60% lower average policy error (APE), 40% faster convergence, 45% reduced runtime complexity, and 40% less sample complexity. Furthermore, M-MEMQ achieves comparable APE with significantly lower complexity than centralized methods. Simulations validate the theoretical analyses.

强化学习无线网络多智能体去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。