用强化学习统一解决无线网络数据新鲜度问题
A Survey of Freshness-Aware Wireless Networking with Reinforcement Learning
- 按更新控制、接入调度等任务分类,构建新鲜度优化的强化学习框架
- 梳理了AoI及其变体在B5G/6G中的建模体系,覆盖三类新鲜度范式
- 适合研究下一代无线网络中智能调度与分布式协同的学者
信息年龄(AoI)已成为现代无线系统中衡量数据新鲜度的核心指标。现有综述或聚焦经典AoI模型,或泛泛讨论强化学习在无线网络中的应用,未将新鲜度作为统一的学习问题来处理。本文从AoI视角出发,系统梳理了原生型、函数型和应用导向型三类新鲜度建模方法,为未来B5G与6G系统中的新鲜度建模提供清晰框架。在此基础上,提出以策略为核心的分类体系,涵盖更新控制强化学习、介质访问强化学习、风险敏感强化学习及多智能体强化学习,全面覆盖采样、调度、轨迹规划、介质访问与分布式协调等关键决策。综述了近年来基于强化学习的实时新鲜度控制进展,并指出延迟决策、随机波动与跨层设计等开放挑战。旨在建立下一代无线网络中基于学习的新鲜度优化统一基础。
原文摘要 · Abstract (English)
The age of information (AoI) has become a central measure of data freshness in modern wireless systems, yet existing surveys either focus on classical AoI formulations or provide broad discussions of reinforcement learning (RL) in wireless networks without addressing freshness as a unified learning problem. Motivated by this gap, this survey examines RL specifically through the lens of AoI and generalized freshness optimization. We organize AoI and its variants into native, function-based, and application-oriented families, providing a clearer view of how freshness should be modeled in B5G and 6G systems. Building on this foundation, we introduce a policy-centric taxonomy that reflects the decisions most relevant to freshness, consisting of update-control RL, medium-access RL, risk-sensitive RL, and multi-agent RL. This structure provides a coherent framework for understanding how learning can support sampling, scheduling, trajectory planning, medium access, and distributed coordination. We further synthesize recent progress in RL-driven freshness control and highlight open challenges related to delayed decision processes, stochastic variability, and cross-layer design. The goal is to establish a unified foundation for learning-based freshness optimization in next-generation wireless networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。