提出新框架,让状态更新更有效应对系统不稳定性。
Beyond Freshness and Semantics: A Coupon-Collector Framework for Effective Status Updates
- 将过期信息建模为可失效的优惠券,设计双阈值调度策略。
- 实验显示比传统方法最高提升50%收益,适应状态紧迫性。
- 无需知道信道或寿命分布,智能学习最优传输时机。
在不可靠、能量受限的无线信道中,针对韦弗长期存在的第C级问题——我的数据包是否真的改善了系统行为?每个新鲜样本都有由系统不稳定性决定的随机失效时间,之后信息对控制无效。将问题建模为带过期优惠券的收集者问题,本文(1)构建二维平均奖励马尔可夫决策过程;(2)证明最优调度策略在接收端新鲜度计时器和发送端存储寿命上均为双阈值;(3)推导出确定性寿命下的闭式策略;(4)设计结构感知强化学习算法(SAQ),无需知晓信道成功率或寿命分布即可学习最优策略。仿真验证理论预测:SAQ性能接近最优值迭代,收敛速度远超基线Q-learning;考虑失效的调度比基于年龄的基线最高提升50%奖励,通过适配状态相关紧急程度,在资源受限下实现第C级有效性。
原文摘要 · Abstract (English)
For status update systems operating over unreliable energy-constrained wireless channels, we address Weaver's long-standing Level-C question: do my packets actually improve the plant's behavior? Each fresh sample carries a stochastic expiration time -- governed by the plant's instability dynamics -- after which the information becomes useless for control. Casting the problem as a coupon-collector variant with expiring coupons, we (i) formulate a two-dimensional average-reward MDP, (ii) prove that the optimal schedule is doubly thresholded in the receiver's freshness timer and the sender's stored lifetime, (iii) derive a closed-form policy for deterministic lifetimes, and (iv) design a Structure-Aware Q-learning algorithm (SAQ) that learns the optimal policy without knowing the channel success probability or lifetime distribution. Simulations validate our theoretical predictions: SAQ matches optimal Value Iteration performance while converging significantly faster than baseline Q-learning, and expiration-aware scheduling achieves up to 50% higher reward than age-based baselines by adapting transmissions to state-dependent urgency -- thereby delivering Level-C effectiveness under tight resource constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。