arXiv:2602.09765cs.RO2026-02被引 8

用视频模型实现零样本3D导航,让AI看懂指令直接规划路径。

NavDreamer: Video Models as Zero-Shot 3D Navigators

  • 用生成式视频模型连接语言指令与导航轨迹,无需训练即可理解动作。
  • 在未见场景中成功导航新物体,跨环境泛化能力显著优于传统方法。
  • 适合研究视觉-语言-动作系统、具身智能及零样本任务部署的开发者。

以往视觉-语言-动作模型在导航任务中面临数据稀缺、采集成本高以及静态表征无法捕捉时空动态与物理规律的问题。本文提出NavDreamer,一种基于视频的3D导航框架,利用生成式视频模型作为语言指令与导航轨迹之间的通用接口。核心假设是:视频兼具时空信息与物理动态表达能力,且互联网规模数据可得,因而具备强大的零样本泛化潜力。为缓解生成预测的随机性,引入基于采样的优化方法,使用视觉语言模型(VLM)对轨迹进行评分与筛选;再通过逆动力学模型从生成视频计划中解码出可执行的路径点。为系统评估该范式在多种视频模型主干上的表现,构建了一个综合性基准,涵盖物体导航、精确导航、空间定位、语言控制与场景推理。大量实验表明,模型在新物体和未见过的环境中均表现出稳健的泛化能力;消融实验揭示,导航的高层决策特性使其特别适合基于视频的规划策略。

原文摘要 · Abstract (English)

Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture temporal dynamics and physical laws. We propose NavDreamer, a video-based framework for 3D navigation that leverages generative video models as a universal interface between language instructions and navigation trajectories. Our main hypothesis is that video's ability to encode spatiotemporal information and physical dynamics, combined with internet-scale availability, enables strong zero-shot generalization in navigation. To mitigate the stochasticity of generative predictions, we introduce a sampling-based optimization method that utilizes a VLM for trajectory scoring and selection. An inverse dynamics model is employed to decode executable waypoints from generated video plans for navigation. To systematically evaluate this paradigm in several video model backbones, we introduce a comprehensive benchmark covering object navigation, precise navigation, spatial grounding, language control, and scene reasoning. Extensive experiments demonstrate robust generalization across novel objects and unseen environments, with ablation studies revealing that navigation's high-level decision-making nature makes it particularly suited for video-based planning.

3D导航视频生成零样本具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。