arXiv:2601.18069cs.NIcs.AI2026-01被引 1

用扩散模型提升信息更新调度的可靠性,兼顾平均延迟与极端风险。

Diffusion Model-based Reinforcement Learning for Version Age of Information Scheduling: Average and Tail-Risk-Sensitive Control

  • 引入扩散机制生成动作,增强策略表达能力
  • 通过分布式评论器实现尾部风险优化,降低极端延迟事件
  • 适合对可靠性要求高的多用户无线系统应用

在实时无线系统中,确保信息及时且语义准确至关重要。传统时效性度量AoI仅反映时间新鲜度,而版本时效性VAoI则考虑发送端与接收端间版本演进,捕捉语义过时。现有方法多聚焦于最小化平均VAoI,忽视了在随机丢包和不可靠信道下罕见但严重的过时事件对可靠性的威胁。本文研究具有长期传输成本约束的多用户状态更新系统中的平均导向与尾部风险敏感的VAoI调度问题。首先将平均VAoI最小化建模为带约束的马尔可夫决策过程,提出基于深度扩散的软演员-评论家(D2SAC)算法,通过扩散去噪过程生成动作,提升策略表达力,建立均值性能基准。在此基础上,提出风险敏感的分布式扩散型软演员-评论家算法RS-D3SAC,融合扩散型演员与分位数式分布评论器,显式建模完整的VAoI回报分布,通过条件风险价值(CVaR)实现严谨的尾部风险优化,同时满足长期传输成本约束。大量仿真表明:尽管D2SAC降低了平均VAoI,RS-D3SAC仍能持续显著降低CVaR,且不牺牲均值性能。尾部风险的主导下降源于分布评论器,扩散演员则提供互补优化,稳定并丰富策略决策,凸显其在多用户无线系统中实现鲁棒、风险感知调度的有效性。

原文摘要 · Abstract (English)

Ensuring timely and semantically accurate information delivery is critical in real-time wireless systems. While Age of Information (AoI) quantifies temporal freshness, Version Age of Information (VAoI) captures semantic staleness by accounting for version evolution between transmitters and receivers. Existing VAoI scheduling approaches primarily focus on minimizing average VAoI, overlooking rare but severe staleness events that can compromise reliability under stochastic packet arrivals and unreliable channels. This paper investigates both average-oriented and tail-risk-sensitive VAoI scheduling in a multi-user status update system with long-term transmission cost constraints. We first formulate the average VAoI minimization problem as a constrained Markov decision process and introduce a deep diffusion-based Soft Actor-Critic (D2SAC) algorithm. By generating actions through a diffusion-based denoising process, D2SAC enhances policy expressiveness and establishes a strong baseline for mean performance. Building on this foundation, we put forth RS-D3SAC, a risk-sensitive deep distributional diffusion-based Soft Actor-Critic algorithm. RS-D3SAC integrates a diffusion-based actor with a quantile-based distributional critic, explicitly modeling the full VAoI return distribution. This enables principled tail-risk optimization via Conditional Value-at-Risk (CVaR) while satisfying long-term transmission cost constraints. Extensive simulations show that, while D2SAC reduces average VAoI, RS-D3SAC consistently achieves substantial reductions in CVaR without sacrificing mean performance. The dominant gain in tail-risk reduction stems from the distributional critic, with the diffusion-based actor providing complementary refinement to stabilize and enrich policy decisions, highlighting their effectiveness for robust and risk-aware VAoI scheduling in multi-user wireless systems.

扩散模型风险控制无线调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。