arXiv:2603.09241cs.CVcs.RO2026-03被引 6

在密集视觉空间建模导航动态,提升路径规划精度

RAE-NWM: Navigation World Model in Dense Visual Representation Space

  • 用DINOv2稠密特征替代压缩潜空间建模动作转移
  • 连续生成中引入时序门控,精准控制动作注入强度
  • 在多个环境上验证,结构稳定性和动作准确率显著提升

视觉导航要求智能体在复杂环境中通过感知与规划抵达目标。世界模型通过模拟动作条件下的状态转移来预测未来观测。当前导航世界模型通常在变分自编码器的压缩潜空间中学习状态演化,但空间压缩常丢失细粒度结构信息,影响精确控制。为更好理解不同表征的传播特性,我们进行线性动力学探测,发现稠密DINOv2特征对动作条件转移具有更强的线性可预测性。受此启发,我们提出基于表示自编码器的导航世界模型(RAE-NWM),在稠密视觉表示空间中建模导航动态。采用带解耦扩散变压器头的条件扩散变压器(CDiT-DH)建模连续转移,并引入独立的时序驱动门控模块,调节生成过程中动作注入强度。大量实验表明,在该空间中建模序列回放可提升结构稳定性与动作准确性,显著改善下游规划与导航性能。

原文摘要 · Abstract (English)

Visual navigation requires agents to reach goals in complex environments through perception and planning. World models address this task by simulating action-conditioned state transitions to predict future observations. Current navigation world models typically learn state evolution under actions within the compressed latent space of a Variational Autoencoder, where spatial compression often discards fine-grained structural information and hinders precise control. To better understand the propagation characteristics of different representations, we conduct a linear dynamics probe and observe that dense DINOv2 features exhibit stronger linear predictability for action-conditioned transitions. Motivated by this observation, we propose the Representation Autoencoder-based Navigation World Model (RAE-NWM), which models navigation dynamics in a dense visual representation space. We employ a Conditional Diffusion Transformer with Decoupled Diffusion Transformer head (CDiT-DH) to model continuous transitions, and introduce a separate time-driven gating module for dynamics conditioning to regulate action injection strength during generation. Extensive evaluations show that modeling sequential rollouts in this space improves structural stability and action accuracy, benefiting downstream planning and navigation.

视觉导航世界模型扩散模型表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。