arXiv:2412.03572cs.CVcs.AI2024-12CVPR被引 305

用视觉动作预测未来画面,让智能体自主规划导航路径。

Navigation World Models

论文配图:Navigation World Models
图 1 · 摘自论文原文
  • 基于条件扩散变换器,从过往观察和动作预测未来视觉
  • 在熟悉环境中可模拟路径并评估是否达成目标,支持动态约束
  • 仅需一张图就能想象陌生环境中的路径,适合新场景导航

导航是具备视觉-运动能力智能体的基本技能。本文提出导航世界模型(NWM),一种可控视频生成模型,可根据历史观察与导航动作预测未来的视觉输入。为捕捉复杂环境动态,NWM采用条件扩散变压器(CDiT),在涵盖人类与机器人视角的多样化第一人称视频数据上训练,参数规模达10亿。在熟悉环境中,NWM可通过模拟轨迹并评估其是否达成目标来规划导航路径。与行为固定的监督导航策略不同,NWM可在规划中动态引入约束。实验表明,该模型能从零开始规划路径,或对来自外部策略采样的路径进行排序。此外,利用学习到的视觉先验,NWM仅需单张输入图像即可在陌生环境中生成想象路径,成为下一代导航系统中灵活且强大的工具。

原文摘要 · Abstract (English)

Navigation is a fundamental skill of agents with visual-motor capabilities. We introduce a Navigation World Model (NWM), a controllable video generation model that predicts future visual observations based on past observations and navigation actions. To capture complex environment dynamics, NWM employs a Conditional Diffusion Transformer (CDiT), trained on a diverse collection of egocentric videos of both human and robotic agents, and scaled up to 1 billion parameters. In familiar environments, NWM can plan navigation trajectories by simulating them and evaluating whether they achieve the desired goal. Unlike supervised navigation policies with fixed behavior, NWM can dynamically incorporate constraints during planning. Experiments demonstrate its effectiveness in planning trajectories from scratch or by ranking trajectories sampled from an external policy. Furthermore, NWM leverages its learned visual priors to imagine trajectories in unfamiliar environments from a single input image, making it a flexible and powerful tool for next-generation navigation systems.

导航建模扩散模型视觉预测路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。