用联合信念优化远程估计,降低信息过时错误度。
Joint Age-State Belief is All You Need: Minimizing AoII via Pull-Based Remote Estimation
- 构建源状态与年龄的联合信念,动态跟踪信息质量。
- 提出基于信念的MAP估计器,相比传统方法更准确。
- 设计深度强化学习与阈值策略,有效最小化AoII。
年龄-状态信念(Age-State Belief)是近期提出的时效性与不匹配度量指标,惩罚错误估计及其持续时间。因此,追踪该指标需同时掌握源过程与估计过程。本文研究在采样率受限下的时隙制拉取式远程估计系统,其中信息源为一般离散时间马尔可夫链(DTMC)。此外,从源到监控端的包传输时间非零,导致监控端无法实时获取实际的AoII过程。为此,我们提出监控端维护一个充分统计量——信念,即从历史观测中获得的状态与年龄联合分布。基于此信念,首先提出一种最大后验(MAP)估计器,替代文献中常见的鞅估计器;其次,通过信念-马尔可夫决策过程(belief-MDP)推导出最优性方程;最后,提出两种依赖信念的策略:一种基于深度强化学习,另一种为基于瞬时期望AoII的阈值策略。
原文摘要 · Abstract (English)
Age of incorrect information (AoII) is a recently proposed freshness and mismatch metric that penalizes an incorrect estimation along with its duration. Therefore, keeping track of AoII requires the knowledge of both the source and estimation processes. In this paper, we consider a time-slotted pull-based remote estimation system under a sampling rate constraint where the information source is a general discrete-time Markov chain (DTMC) process. Moreover, packet transmission times from the source to the monitor are non-zero which disallows the monitor to have perfect information on the actual AoII process at any time. Hence, for this pull-based system, we propose the monitor to maintain a sufficient statistic called {\em belief} which stands for the joint distribution of the age and source processes to be obtained from the history of all observations. Using belief, we first propose a maximum a posteriori (MAP) estimator to be used at the monitor as opposed to existing martingale estimators in the literature. Second, we obtain the optimality equations from the belief-MDP (Markov decision process) formulation. Finally, we propose two belief-dependent policies one of which is based on deep reinforcement learning, and the other one is a threshold-based policy based on the instantaneous expected AoII.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。