arXiv:2505.02640cs.LGcs.AI2025-05被引 1

为动态资源受限的物联网设计自适应预算强化学习算法

Adaptive Budgeted Multi-Armed Bandits for IoT with Dynamic Resource Constraints

  • 引入衰减违规预算,允许早期有限越界,逐步强化约束遵守
  • 理论证明该算法在学习期内实现亚线性损失和对数级违规次数
  • 适合需要实时响应与资源自适应的物联网系统部署

物联网(IoT)系统日益运行在设备需实时响应且资源约束(如能量、带宽)动态变化的环境中。现有方法难以应对随时间演化的操作限制。为此,我们提出一种专为物联网设计的新型有预算多臂老虎机框架。模型引入衰减违规预算机制,允许学习初期有限度违反约束,并随时间推移逐步加强合规性。我们提出了有预算上置信界(Budgeted UCB)算法,自适应平衡性能优化与动态约束遵守。理论分析表明,该算法在学习周期内达到亚线性后悔率和对数级约束违规次数。在无线通信场景中的大量仿真显示,该方法比标准在线学习方法具有更快的适应速度和更优的约束满足能力。结果凸显了该框架在构建自适应、资源感知型物联网系统方面的潜力。

原文摘要 · Abstract (English)

Internet of Things (IoT) systems increasingly operate in environments where devices must respond in real time while managing fluctuating resource constraints, including energy and bandwidth. Yet, current approaches often fall short in addressing scenarios where operational constraints evolve over time. To address these limitations, we propose a novel Budgeted Multi-Armed Bandit framework tailored for IoT applications with dynamic operational limits. Our model introduces a decaying violation budget, which permits limited constraint violations early in the learning process and gradually enforces stricter compliance over time. We present the Budgeted Upper Confidence Bound (UCB) algorithm, which adaptively balances performance optimization and compliance with time-varying constraints. We provide theoretical guarantees showing that Budgeted UCB achieves sublinear regret and logarithmic constraint violations over the learning horizon. Extensive simulations in a wireless communication setting show that our approach achieves faster adaptation and better constraint satisfaction than standard online learning methods. These results highlight the framework's potential for building adaptive, resource-aware IoT systems.

物联网强化学习动态约束多臂老虎机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。