用世界模型模拟车联网络,高效优化长期信息时效性。
World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks
- 构建环境世界模型,通过想象轨迹学习长期调度策略
- 在毫米波车联网中实现26%和16%的年龄信息改善
- 适合高动态、低反馈频次的智能交通系统研究
传统基于强化学习的无线网络学习方法依赖昂贵的试错机制和实时反馈,导致数据效率低且策略短视。在高不确定性、需长期规划的复杂动态网络中尤为突出。本文提出一种基于世界模型的学习框架,以最小化车辆网络中的包完整感知年龄信息(CAoI)。针对毫米波车联通信(mmWave V2X)这一具有高移动性、频繁信号遮挡和极短相干时间的典型场景,构建联合学习动态环境模型与生成想象轨迹的框架。长期策略在可微分的想象轨迹中学习,而非真实环境交互。由于具备想象力,该模型可在无实际观测时持续做出高效决策。在基于Sionna的物理层端到端信道建模仿真器上进行大量实验,集成射线追踪与场景几何及材料属性。结果表明,所提方法在数据效率上显著提升,在CAoI上分别优于基于模型的强化学习(MBRL)和无模型强化学习(MFRL)方法26%和16%。
原文摘要 · Abstract (English)
Traditional reinforcement learning (RL)-based learning approaches for wireless networks rely on expensive trial-and-error mechanisms and real-time feedback based on extensive environment interactions, which leads to low data efficiency and short-sighted policies. These limitations become particularly problematic in complex, dynamic networks with high uncertainty and long-term planning requirements. To address these limitations, in this paper, a novel world model-based learning framework is proposed to minimize packet-completeness-aware age of information (CAoI) in a vehicular network. Particularly, a challenging representative scenario is considered pertaining to a millimeter-wave (mmWave) vehicle-to-everything (V2X) communication network, which is characterized by high mobility, frequent signal blockages, and extremely short coherence time. Then, a world model framework is proposed to jointly learn a dynamic model of the mmWave V2X environment and use it to imagine trajectories for learning how to perform link scheduling. In particular, the long-term policy is learned in differentiable imagined trajectories instead of environment interactions. Moreover, owing to its imagination abilities, the world model can jointly predict time-varying wireless data and optimize link scheduling in real-world wireless and V2X networks. Thus, during intervals without actual observations, the world model remains capable of making efficient decisions. Extensive experiments are performed on a realistic simulator based on Sionna that integrates physics-based end-to-end channel modeling, ray-tracing, and scene geometries with material properties. Simulation results show that the proposed world model achieves a significant improvement in data efficiency, and achieves 26% improvement and 16% improvement in CAoI, respectively, compared to the model-based RL (MBRL) method and the model-free RL (MFRL) method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。