用强化学习优化物联网信息更新,让数据更实时且省成本。
Policy Gradient Algorithms for Age-of-Information Cost Minimization
- 基于无模型强化学习,用策略梯度在线优化更新时机。
- 平均代价接近最优值(误差小于3%),且计算开销低一个数量级。
- 适用于多种场景,可并行使用两种策略进一步降本。
近年来,网络物理系统的发展使得提升环境信息的新鲜度变得愈发重要。然而,在传输延迟和年龄代价函数未知的情况下,优化物联网设备的信息访问策略以最大化数据新鲜度(以年龄-信息度量,AoI 衡量)是一项挑战。本文提出两种算法,针对可即时生成(generate-at-will)模型下的网络物理系统,通过在线策略优化信息更新过程,目标是最小化时间平均代价,该代价整合了接收端的 AoI 和数据传输成本,适用性广。两类算法均在无模型强化学习框架下采用策略梯度方法,专为连续状态与动作空间设计。每种算法采用不同的更新决策策略,且可同时使用,带来额外成本降低。实验表明,所提算法具备良好收敛性,在最优值可计算时,时间平均代价仅比最优值高3%以下;相较于现有先进方法,其在适用场景范围、更低的时间平均代价以及至少低一个数量级的计算开销方面表现更优。
原文摘要 · Abstract (English)
Recent developments in cyber-physical systems have increased the importance of maximizing the freshness of the information about the physical environment. However, optimizing the access policies of Internet of Things devices to maximize the data freshness, measured as a function of the Age-of-Information (AoI) metric, is a challenging task. This work introduces two algorithms to optimize the information update process in cyber-physical systems operating under the generate-at-will model, by finding an online policy without knowing the characteristics of the transmission delay or the age cost function. The optimization seeks to minimize the time-average cost, which integrates the AoI at the receiver and the data transmission cost, making the approach suitable for a broad range of scenarios. Both algorithms employ policy gradient methods within the framework of model-free reinforcement learning (RL) and are specifically designed to handle continuous state and action spaces. Each algorithm minimizes the cost using a distinct strategy for deciding when to send an information update. Moreover, we demonstrate that it is feasible to apply the two strategies simultaneously, leading to an additional reduction in cost. The results demonstrate that the proposed algorithms exhibit good convergence properties and achieve a time-average cost within 3% of the optimal value, when the latter is computable. A comparison with other state-of-the-art methods shows that the proposed algorithms outperform them in one or more of the following aspects: being applicable to a broader range of scenarios, achieving a lower time-average cost, and requiring a computational cost at least one order of magnitude lower.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。