arXiv:2601.05956cs.LG2026-01

用包龄替代队列长度,让无线调度更抗环境突变

On the Robustness of Age for Learning-Based Wireless Scheduling in Unknown Environments

  • 用头包年龄代替虚拟队列长度做学习调度
  • 在独立同分布环境下性能达当前最优
  • 突变信道下仍稳定,约束失效后快速恢复

约束性组合多臂赌博机模型广泛应用于无线网络中的调度问题,尤其在未知信道条件下优化吞吐量。现有方法通常将多臂赌博机与虚拟队列技术结合,通过最小化虚拟队列长度来追踪吞吐量约束违反情况。然而,在信道条件突然变化时,此类约束可能不可行,导致虚拟队列长度无界增长。本文提出关键观察:虚拟队列中最早到达包的年龄(即头包年龄)在算法设计中更具鲁棒性。为此,我们设计了一种基于学习的调度策略,以头包年龄替代虚拟队列长度。实验表明,该策略在独立同分布网络条件下达到现有最优性能;更重要的是,在信道条件突变时系统仍保持稳定,并能快速从约束不可行期恢复。

原文摘要 · Abstract (English)

The constrained combinatorial multi-armed bandit model has been widely employed to solve problems in wireless networking and related areas, including the problem of wireless scheduling for throughput optimization under unknown channel conditions. Most work in this area uses an algorithm design strategy that combines a bandit learning algorithm with the virtual queue technique to track the throughput constraint violation. These algorithms seek to minimize the virtual queue length in their algorithm design. However, in networks where channel conditions change abruptly, the resulting constraints may become infeasible, leading to unbounded growth in virtual queue lengths. In this paper, we make the key observation that the dynamics of the head-of-line age, i.e. the age of the oldest packet in the virtual queue, make it more robust when used in algorithm design compared to the virtual queue length. We therefore design a learning-based scheduling policy that uses the head-of-line age in place of the virtual queue length. We show that our policy matches state-of-the-art performance under i.i.d. network conditions. Crucially, we also show that the system remains stable even under abrupt changes in channel conditions and can rapidly recover from periods of constraint infeasibility.

无线调度强化学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。