让机器人直接用视觉预测行动,无需额外规划器。
NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

- 用扩散-变换器联合学习未来视觉、目标进展与动作序列。
- 在离线评测和真实机器人部署中均超越基于规划的模型。
- 适合需要闭环控制的视觉导航任务,尤其注重效率的场景。
目标条件视觉导航要求机器人在部分可观测环境下,通过预判自身运动如何改变当前视角,并判断该变化是否使其更接近目标。导航世界模型可提供此类视觉前瞻性,但仍是仅能预测的模块,需依赖外部规划器将预测结果转化为闭环控制。本文提出导航世界动作模型(NavWAM),一种扩散-变换器策略,通过在共享潜在序列中表示未来观测、目标进展值与动作块,将导航世界模型的预测直接转化为可执行动作。通过联合学习未来预测与决定闭环行为的动作及价值目标,NavWAM使视觉前瞻性可直接用于机器人控制。我们通过仿真预训练与真实机器人适应构建了NavWAM,对比基于规划的世界模型和代表性直接导航策略,在图像目标导航任务上进行评估。无论在离线基准测试还是闭环真实机器人部署中,NavWAM均优于基于规划的世界模型基线,且未使用CEM风格的动作搜索,仅以默认策略模式运行。
原文摘要 · Abstract (English)
Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change brings it closer to the goal. Navigation world models provide such visual foresight, but they remain prediction modules that require an external planner to convert predicted futures into closed-loop control. We propose Navigation World Action Model (NavWAM), a diffusion-transformer policy that turns navigation world-model prediction into executable action by representing future observations, goal-progress values, and action chunks in a shared latent sequence. By learning future prediction jointly with the action and value targets that determine closed-loop behavior, NavWAM makes visual foresight directly usable for robot control. We build NavWAM through simulation pretraining and real-robot adaptation, and evaluate it on image-goal navigation against planning-based world models and a representative direct navigation policy. Across offline benchmarks and closed-loop real-robot deployment, NavWAM improves over planning-based world-model baselines in our evaluations while using the default policy mode without CEM-style action search. Project page: https://dachii-azm.github.io/navwam/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。