arXiv:2506.18186cs.LGstat.ML2025-06被引 3

提出在线自适应的滑动窗口策略,解决非平稳环境下的资源分配难题。

Online Learning of Whittle Indices for Restless Bandits with Non-Stationary Transition Kernels

  • 基于滑动窗口和置信上界估计动态变化的转移概率
  • 理论证明动态损失增长低于线性,适应能力更强
  • 适合动态网络、物联网等时变系统中的实时决策

随机多臂老虎机(RMAB)是解决网络化系统资源分配问题的常用框架。本文研究在未知且非平稳动态环境下最优资源分配问题。即使已知完整模型参数,求解RMAB也被证明为PSPACE难问题。尽管惠特尔指数策略具有渐近最优性和低计算开销,但其依赖于平稳转移核,这在许多现代网络应用中不现实。为此,我们提出一种滑动窗口在线惠特尔(SW-Whittle)策略,在保持计算高效的同时适应随时间变化的转移核。理论分析表明,该算法在轮次数量上实现了次线性动态遗憾。针对变差预算事先未知的情况,我们将带间带(Bandit-over-Bandit)框架与滑动窗口设计结合:窗口长度在线根据估计的变差调整,惠特尔指数通过转移核的置信上界和双线性优化计算。数值实验显示,该算法在多种非平稳环境中持续优于基线方法,累积遗憾最低。

原文摘要 · Abstract (English)

The restless multi-armed bandit (RMAB) framework is a popular approach to solving resource allocation problems in networked systems. In this paper, we study optimal resource allocation in RMABs facing unknown and non-stationary dynamics. Solving RMABs optimally is known to be PSPACE-hard even with full knowledge of model parameters. While Whittle index policies offer asymptotic optimality with low computational cost, they require access to stationary transition kernels, an unrealistic assumption in many modern networking applications. To address this challenge, we propose a Sliding-Window Online Whittle (SW-Whittle) policy that remains computationally efficient while adapting to time-varying kernels. Through theoretical analysis, we show that our algorithm achieves sub-linear dynamic regret with respect to the number of episodes. We further address the important case where the variation budget is unknown in advance by combining a Bandit-over-Bandit framework with our sliding-window design. In our scheme, window lengths are tuned online as a function of the estimated variation, while Whittle indices are computed via an upper-confidence-bound of the estimated transition kernels and a bilinear optimization routine. Numerical experiments demonstrate that our algorithm consistently outperforms baselines, achieving the lowest cumulative regret across a range of non-stationary environments.

强化学习在线学习资源分配非平稳系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。