让机器人看懂未来,提前规划路线,导航更稳更准。
WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

- 用统一模型同时生成动作和视觉预测,避免逐步推演的延迟与误差
- 在复杂场景中提升15.7%图像目标导航成功率,3.3%点目标导航成功率
- 支持多种目标类型,可直接从仿真环境零样本迁移到真实世界
视觉导航需在复杂几何与物理约束下生成平滑无碰撞轨迹。现有反应式策略直接将观测映射为动作,缺乏前瞻性,难以主动避障;虽有视觉想象可提供预测能力,但传统模块化方法将场景预测与策略学习分离,易导致误差累积和推理效率低。为此,我们提出WAM-Nav:一种用于具身视觉导航的隐空间世界-动作联合建模方法,通过共享扩散变压器实现异构联合扩散,同步生成长时程动作与短时程视觉预见,降低多步自回归推演带来的推理延迟与视觉误差积累。为进一步促进轨迹平滑一致,引入双流上下文条件机制,融合全局自身运动历史与序列视觉观测;结合统一目标对齐模块,保持不同目标类型间的表征平衡,使单个策略自然支持图像目标、点目标及无目标探索。在挑战性数据集ClutterScenes与InternScenes上的实验表明,WAM-Nav具有强泛化能力,尤其在图像目标与点目标导航上分别提升成功率达15.7%和3.3%。真实世界部署验证了有效的零样本仿真到现实迁移,在多样室内室外环境中平均任务成功率达85%。
原文摘要 · Abstract (English)
Visual navigation requires generating smooth and collision-free trajectories under complex geometric and physical constraints. Existing reactive policies that directly map observations to actions lack anticipatory reasoning, limiting their ability to proactively avoid obstacles. While visual imagination offers predictive foresight, conventional modular approaches separate scene prediction from policy learning, often leading to error accumulation and inefficient inference. To address these limitations, we propose WAM-Nav, a Latent World-Action Model for embodied visual navigation that jointly learns action generation and latent visual foresight, enabling more robust and foresighted navigation decisions without compromising inference efficiency. Specifically, WAM-Nav utilizes a shared Diffusion Transformer for asymmetric joint diffusion to concurrently generate long-horizon actions and short-horizon visual foresight, reducing the inference latency and visual error accumulation inherent in multi-step autoregressive rollouts. To further encourage smooth and consistent trajectory generation, we introduce a dual-stream contextual conditioning mechanism that integrates episode-level ego-motion history with sequential visual observations. Combined with a unified goal alignment module that preserves balanced representations across goal types, WAM-Nav naturally supports Image-Goal, Point-Goal, and No-Goal exploration within a single policy. Extensive experiments on the challenging ClutterScenes and InternScenes benchmarks demonstrate strong generalization of WAM-Nav, particularly on Image-Goal and Point-Goal navigation, where it improves success rates by 15.7% and 3.3%, respectively. Real-world deployment further validates effective zero-shot sim-to-real transfer, achieving an average 85% task success rate across diverse indoor and outdoor environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。