arXiv:2510.24546cs.ITcs.LG2025-10被引 7

提出双脑世界模型,让无线网络学会长期规划与自适应决策。

Dual-Mind World Models: A General Framework for Learning in Dynamic Wireless Networks

  • 用双系统模拟认知:一个学统计模式,一个推逻辑规律。
  • 在毫米波车联网中实现数据效率提升与未见环境泛化能力。
  • 无需实时观测也能持续做决策,适合高动态复杂网络场景。

尽管强化学习在无线网络中广泛应用,但现有基于无模型强化学习(MFRL)和基于模型强化学习(MBRL)的方法存在数据效率低、视野短的缺陷。这些方法仅捕捉统计规律,无法理解无线数据背后的物理与逻辑,难以推广至新网络状态。针对复杂高动态无线网络中的长期规划需求,本文提出一种新型双脑世界模型学习框架,旨在优化毫米波车联网场景下的完备加权信息年龄(CAoI)。受认知心理学启发,该框架包含基于模式的System 1与基于逻辑的System 2组件,分别学习网络动态与内在逻辑,并通过端到端可微的想象轨迹实现长期链路调度。调度不依赖环境交互数据,而基于逻辑一致的扩展时域想象。通过想象回溯,模型可联合推理网络状态并规划链路调度。在无观测间隔期间仍能高效决策。实验基于支持真实物理信道、射线追踪与材料属性的Sionna仿真器进行,结果表明,该框架显著提升数据效率,在未见环境中表现强泛化性与适应性,优于当前最优强化学习基线及仅含System 1的模型。

原文摘要 · Abstract (English)

Despite the popularity of reinforcement learning (RL) in wireless networks, existing approaches that rely on model-free RL (MFRL) and model-based RL (MBRL) are data inefficient and short-sighted. Such RL-based solutions cannot generalize to novel network states since they capture only statistical patterns rather than the underlying physics and logic from wireless data. These limitations become particularly challenging in complex wireless networks with high dynamics and long-term planning requirements. To address these limitations, in this paper, a novel dual-mind world model-based learning framework is proposed with the goal of optimizing completeness-weighted age of information (CAoI) in a challenging mmWave V2X scenario. Inspired by cognitive psychology, the proposed dual-mind world model encompasses a pattern-driven System 1 component and a logic-driven System 2 component to learn dynamics and logic of the wireless network, and to provide long-term link scheduling over reliable imagined trajectories. Link scheduling is learned through end-to-end differentiable imagined trajectories with logical consistency over an extended horizon rather than relying on wireless data obtained from environment interactions. Moreover, through imagination rollouts, the proposed world model can jointly reason network states and plan link scheduling. During intervals without observations, the proposed method remains capable of making efficient decisions. Extensive experiments are conducted on a realistic simulator based on Sionna with real-world physical channel, ray-tracing, and scene objects with material properties. Simulation results show that the proposed world model achieves a significant improvement in data efficiency and achieves strong generalization and adaptation to unseen environments, compared to the state-of-the-art RL baselines, and the world model approach with only System 1.

强化学习无线网络世界模型长时序规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。