用扩散模型+平均场通信,让多个无线节点自主高效分配资源。
Multi-Agent Conditional Diffusion Model with Mean Field Communication as Wireless Resource Allocation Planner
- 每个节点用扩散模型预测环境变化,结合逆动力学生成动作。
- 在真实无线网络中平均收益提升32%,服务质量指标更优。
- 适合大规模分布式通信系统,通信开销低且能稳定协作。
在无线通信系统中,高效自适应的资源分配对提升整体服务质量(QoS)至关重要。相比传统无模型强化学习(MFRL),基于模型的强化学习(MBRL)先学习生成世界模型以支持后续规划,重用历史经验可带来更稳定的训练行为。然而,其在大规模无线网络中的部署仍面临高维随机动态、强代理间协作和通信约束等挑战。为此,我们提出多智能体条件扩散模型规划器(MA-CDMP),用于去中心化通信资源管理。基于分布式训练、去中心化执行(DTDE)范式,MA-CDMP将每个通信节点视为独立智能体,采用扩散模型(DMs)捕捉并预测环境动态;同时引入逆动力学模型引导动作生成,提升样本效率与策略可扩展性。为近似大规模智能体交互,设计了平均场(MF)机制辅助扩散模型中的分类器,有效缓解智能体非平稳性,在分布式场景中实现最小通信开销下的协同优化。我们进一步从理论上建立了基于平均场扩散生成的分布近似误差上界,保证了多智能体随机动态建模的收敛稳定性与可靠性。大量实验表明,MA-CDMP在平均奖励和QoS指标上持续优于现有MARL基线,展现出在真实无线网络优化中的可扩展性与实用性。
原文摘要 · Abstract (English)
In wireless communication systems, efficient and adaptive resource allocation plays a crucial role in enhancing overall Quality of Service (QoS). Compared to the conventional Model-Free Reinforcement Learning (MFRL) scheme, Model-Based RL (MBRL) first learns a generative world model for subsequent planning. The reuse of historical experience in MBRL promises more stable training behavior, yet its deployment in large-scale wireless networks remains challenging due to high-dimensional stochastic dynamics, strong inter-agent cooperation, and communication constraints. To overcome these challenges, we propose the Multi-Agent Conditional Diffusion Model Planner (MA-CDMP) for decentralized communication resource management. Built upon the Distributed Training with Decentralized Execution (DTDE) paradigm, MA-CDMP models each communication node as an autonomous agent and employs Diffusion Models (DMs) to capture and predict environment dynamics. Meanwhile, an inverse dynamics model guides action generation, thereby enhancing sample efficiency and policy scalability. Moreover, to approximate large-scale agent interactions, a Mean-Field (MF) mechanism is introduced as an assistance to the classifier in DMs. This design mitigates inter-agent non-stationarity and enhances cooperation with minimal communication overhead in distributed settings. We further theoretically establish an upper bound on the distributional approximation error introduced by the MF-based diffusion generation, guaranteeing convergence stability and reliable modeling of multi-agent stochastic dynamics. Extensive experiments demonstrate that MA-CDMP consistently outperforms existing MARL baselines in terms of average reward and QoS metrics, showcasing its scalability and practicality for real-world wireless network optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。