arXiv:2603.29165cs.CVcs.AI2026-03

让AI提前‘做梦’预测动作后果,提升导航鲁棒性

LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning

  • 用潜空间令牌模拟未来视觉变化,实现无监督前瞻推理
  • 在R2R-CE等3个基准上刷新最佳性能,实机测试表现优异
  • 适合需要理解环境因果关系的智能机器人导航任务

现有视觉语言导航(VLN)模型主要依赖过去和当前视觉观察,忽略动作带来的未来视觉动态。这导致对动作与环境变化间因果关系的理解不足,限制了决策鲁棒性。人类则能通过动作-动态因果性预想近未来,从而提升环境理解和导航选择。受此启发,我们提出LatentPilot,一种新范式:在训练中利用未来观测作为数据源,学习动作条件下的视觉动态,推理时无需访问未来帧。具体地,采用飞轮式训练机制,迭代收集在线轨迹并重训练模型以匹配代理行为分布,当代理偏离过大时触发专家接管。LatentPilot还无监督学习视觉潜变量令牌;这些令牌在连续潜空间中全局注意力,并跨步骤传递,既作为当前输出又作为下一输入,使代理能够‘做梦’并推理动作如何影响后续观察。在R2R-CE、RxR-CE和R2R-PE基准上取得新最优结果,真实机器人在多样环境中测试也验证了其对环境-动作动态的优越理解能力。

原文摘要 · Abstract (English)

Existing vision-and-language navigation (VLN) models primarily reason over past and current visual observations, while largely ignoring the future visual dynamics induced by actions. As a result, they often lack an effective understanding of the causal relationship between actions and how the visual world changes, limiting robust decision-making. Humans, in contrast, can imagine the near future by leveraging action-dynamics causality, which improves both environmental understanding and navigation choices. Inspired by this capability, we propose LatentPilot, a new paradigm that exploits future observations during training as a valuable data source to learn action-conditioned visual dynamics, while requiring no access to future frames at inference. Concretely, we propose a flywheel-style training mechanism that iteratively collects on-policy trajectories and retrains the model to better match the agent's behavior distribution, with an expert takeover triggered when the agent deviates excessively. LatentPilot further learns visual latent tokens without explicit supervision; these latent tokens attend globally in a continuous latent space and are carried across steps, serving as both the current output and the next input, thereby enabling the agent to dream ahead and reason about how actions will affect subsequent observations. Experiments on R2R-CE, RxR-CE, and R2R-PE benchmarks achieve new SOTA results, and real-robot tests across diverse environments demonstrate LatentPilot's superior understanding of environment-action dynamics in scene. Project page:https://abdd.top/latentpilot/

视觉导航前瞻推理潜空间机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。