解决密集网络中能耗与切换的动态平衡问题,实现稳定高效资源管理。
Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints

- 引入虚拟队列的漂移-惩罚分解,将约束转化为每步奖励
- 在有限时域内保持吞吐量与公平性的持续平衡,不提前耗尽预算
- 适合需要长期稳定服务的无线网络场景,如5G/6G基站调度
在密集无线网络中,最大化吞吐量并保证比例公平性需联合管理用户关联、调度、基站激活与切换控制,受限于硬性有限时域内的能量和切换预算,导致基站侧节能与用户侧切换之间的根本矛盾。多智能体强化学习(MARL)虽适合作分布式序列控制框架,但面临两大挑战:有限时域约束无法在每个时间步评估,且非线性比例公平效用无法进行逐时隙合理分解。本文提出HeLyMARL——一种嵌入李雅普诺夫的异构MARL框架,通过虚拟队列实现漂移-惩罚分解。将能量与切换约束压力直接内化为统一的每步奖励,把有约束的有限时域问题转化为无约束MARL问题。对比两种基于拉格朗日的替代方法发现,拉格朗日松弛仅在训练周期间调节约束,而HeLyMARL的虚拟队列能在单个训练周期内的任意部分时域内约束累积预算消耗,实现超越贪婪李雅普诺夫控制的节奏保障。仿真表明,只有HeLyMARL能维持整个时域内的吞吐量-公平性平衡与连续服务,优于传统MARL、李雅普诺夫及约束型MARL基准,且无提前预算耗尽现象。
原文摘要 · Abstract (English)
Maximizing throughput under proportional fairness in dense wireless networks requires jointly managing user association, scheduling, base station (BS) activation, and handover control under hard finite-horizon energy and handover budgets, which induces a fundamental tension between BS-side energy management and user-side handover regulation. While multi-agent reinforcement learning (MARL) is a natural framework for such distributed sequential control, its application here faces two difficulties: finite-horizon budget constraints cannot be evaluated at each time slot, and the nonlinear proportional fairness utility admits no principled per-slot decomposition. We propose HeLyMARL, a Lyapunov-embedded heterogeneous MARL framework that resolves both via drift-plus-penalty decomposition with virtual queues. The energy and handover constraint pressures are internalized directly into a unified per-slot reward, converting the constrained finite-horizon problem into an unconstrained MARL problem. Comparison against two Lagrangian-based alternatives reveals a timescale separation: Lagrangian relaxation regulates constraints only across training episodes, whereas the virtual queues of HeLyMARL bound cumulative budget consumption at every partial horizon within an episode, a pacing guarantee beyond the reach of greedy Lyapunov-based control. Simulations show that HeLyMARL is the only method that sustains the throughput-fairness balance together with uninterrupted service throughout the horizon, outperforming conventional MARL, Lyapunov-based, and constrained MARL benchmarks without premature budget exhaustion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。