arXiv:2601.08760cs.LGcs.MA2026-01

针对边缘网络中信息更新延迟问题,提出自适应请求算法提升数据新鲜度。

Adaptive Requesting in Decentralized Edge Networks via Non-Stationary Bandits

  • 用自适应窗口+周期监测追踪动态变化的请求收益。
  • 在非平稳环境下实现接近最优的请求决策性能。
  • 适合时敏型边缘计算场景,如自动驾驶、工业物联网。

我们研究一种去中心化协同请求问题,旨在优化由多个客户端、接入节点(ANs)和服务器组成的边缘网络中时敏客户端的信息新鲜度。客户端通过充当网关的接入节点请求内容,但无法观测接入节点状态或其它客户端的行为。我们将奖励定义为客户端选择某接入节点所带来的时间信息年龄减少量,并将问题建模为非平稳多臂赌博机。在此去中心化且部分可观测设置下,奖励过程具有历史依赖性且跨客户端耦合,预期奖励同时表现出突变与渐变特征,使经典赌博机方法失效。为此,我们提出自适应重置的AGING BANDIT算法,结合自适应窗口与周期性监控以追踪演化中的奖励分布。我们建立了理论性能保证,证明该算法可达到近似最优性能,并通过仿真验证了理论结果。

原文摘要 · Abstract (English)

We study a decentralized collaborative requesting problem that aims to optimize the information freshness of time-sensitive clients in edge networks consisting of multiple clients, access nodes (ANs), and servers. Clients request content through ANs acting as gateways, without observing AN states or the actions of other clients. We define the reward as the age of information reduction resulting from a client's selection of an AN, and formulate the problem as a non-stationary multi-armed bandit. In this decentralized and partially observable setting, the resulting reward process is history-dependent and coupled across clients, and exhibits both abrupt and gradual changes in expected rewards, rendering classical bandit-based approaches ineffective. To address these challenges, we propose the AGING BANDIT WITH ADAPTIVE RESET algorithm, which combines adaptive windowing with periodic monitoring to track evolving reward distributions. We establish theoretical performance guarantees showing that the proposed algorithm achieves near-optimal performance, and we validate the theoretical results through simulations.

边缘计算信息新鲜度强化学习非平稳

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。