用异构强化学习提升无线网络信道访问效率
Heterogeneous Multi-Agent Reinforcement Learning for Distributed Channel Access in WLANs
- 异构智能体混合使用价值与策略型算法,通过集中训练分散执行协同
- 相比传统机制,吞吐量提升且延迟、碰撞率显著降低
- 适合复杂场景下需兼容旧协议的智能网络部署
本文研究多智能体强化学习(MARL)在无线局域网分布式信道接入中的应用。针对实际中智能体采用不同学习算法(基于价值或策略)的挑战,提出名为QPMIX的异构MARL训练框架,采用集中训练、分散执行范式,实现异构智能体协作。理论上证明了在线性价值函数近似下该方法的收敛性。所提方法最大化网络吞吐量并保障终端公平性,显著提升整体性能。仿真结果表明,在饱和流量场景下,QPMIX相较传统载波侦听多路访问避免冲突(CSMA/CA)机制,有效提升吞吐量,降低平均延迟、延迟抖动和碰撞率;在非饱和及对时延敏感场景下仍具鲁棒性,且能与传统机制共存,促进异构智能体间合作。
原文摘要 · Abstract (English)
This paper investigates the use of multi-agent reinforcement learning (MARL) to address distributed channel access in wireless local area networks. In particular, we consider the challenging yet more practical case where the agents heterogeneously adopt value-based or policy-based reinforcement learning algorithms to train the model. We propose a heterogeneous MARL training framework, named QPMIX, which adopts a centralized training with distributed execution paradigm to enable heterogeneous agents to collaborate. Moreover, we theoretically prove the convergence of the proposed heterogeneous MARL method when using the linear value function approximation. Our method maximizes the network throughput and ensures fairness among stations, therefore, enhancing the overall network performance. Simulation results demonstrate that the proposed QPMIX algorithm improves throughput, mean delay, delay jitter, and collision rates compared with conventional carrier-sense multiple access with collision avoidance (CSMA/CA) mechanism in the saturated traffic scenario. Furthermore, the QPMIX algorithm is robust in unsaturated and delay-sensitive traffic scenarios. It coexists well with the conventional CSMA/CA mechanism and promotes cooperation among heterogeneous agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。