arXiv:2604.07411cs.LGcs.AI2026-04

用奖励机器强化学习优化移动网络睡眠策略,兼顾省电与服务质量。

Reinforcement Learning with Reward Machines for Sleep Control in Mobile Networks

  • 引入奖励机器追踪历史服务质量违规,解决长期约束依赖问题。
  • 在多种流量模式下,实现平均丢包率与最小吞吐量的稳定控制。
  • 适合关注绿色通信与智能资源调度的研究者与工程师。

移动网络能源效率对可持续通信基础设施至关重要,尤其在持续网络密集化导致功耗上升的背景下。通过使网络组件进入休眠状态可降低能耗,但如何在保障服务质量(QoS)的前提下决定哪些组件休眠、何时休眠及休眠时长,仍是复杂优化难题。本文提出基于奖励机器(Reward Machines, RMs)的强化学习方法,用于平衡即时节能与长期QoS影响——即针对截止期限敏感流量的平均丢包率,以及恒定速率用户的时间平均最小吞吐量保证。难点在于,这些时间平均约束依赖于累积性能而非瞬时表现,导致有效奖励具有非马尔可夫性,最优动作需依赖操作历史而非当前系统状态。为此,奖励机器通过维护抽象状态显式追踪随时间累积的QoS违规情况。本框架为下一代移动网络在多样化流量模式和QoS需求下的能源管理提供了原理严谨且可扩展的解决方案。

原文摘要 · Abstract (English)

Energy efficiency in mobile networks is crucial for sustainable telecommunications infrastructure, particularly as network densification continues to increase power consumption. Sleep mechanisms for the components in mobile networks can reduce energy use, but deciding which components to put to sleep, when, and for how long while preserving quality of service (QoS) remains a difficult optimisation problem. In this paper, we utilise reinforcement learning with reward machines (RMs) to make sleep-control decisions that balance immediate energy savings and long-term QoS impact, i.e. time-averaged packet drop rates for deadline-constrained traffic and time-averaged minimum-throughput guarantees for constant-rate users. A challenge is that time-averaged constraints depend on cumulative performance over time rather than immediate performance. As a result, the effective reward is non-Markovian, and optimal actions depend on operational history rather than the instantaneous system state. RMs account for the history dependence by maintaining an abstract state that explicitly tracks the QoS constraint violations over time. Our framework provides a principled, scalable approach to energy management for next-generation mobile networks under diverse traffic patterns and QoS requirements.

强化学习节能移动网络服务质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。