arXiv:2604.07957cs.AIcs.CV2026-04

用生成世界模型构建结构化监督,提升视觉语言导航轨迹预测精度。

WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models

  • 通过生成未来视图构建语义空间记忆,生成导航伪标签。
  • 在Target-Bench上将ADE降低18.0%,FDE降低42.1%。
  • 适合研究视觉语言导航与世界模型融合的学者参考。

视觉语言模型(VLMs)和生成式世界模型为具身导航带来了新机遇。尽管当前的VLM常生成不稳定的轨迹,而世界模型虽能合成合理未来画面,却无法直接提供导航学习所需的锚定信号。本文提出WorldMAP框架,将世界模型生成的未来画面转化为持久的语义-空间结构与规划衍生的监督信号。其教师模型基于生成视频构建语义空间记忆,定位任务相关目标与障碍,并通过显式规划生成轨迹伪标签。学生模型则采用多假设轨迹头,直接从视觉-语言输入中预测导航轨迹。在Target-Bench上,WorldMAP优于对比方法,相对最佳基线分别将ADE降低18.0%、FDE降低42.1%;同时使小型开源VLM的DTW性能达到与专有模型相当水平。结果表明,在具身导航中,世界模型的价值可能不在于提供可执行的想象证据,而在于生成结构化监督信号以支持导航学习。

原文摘要 · Abstract (English)

Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct planners or trajectory predictors, while world models support look-ahead reasoning by imagining future views. Yet predicting a reliable trajectory from a single egocentric observation remains challenging. Current VLMs often generate unstable trajectories, and world models, though able to synthesize plausible futures, do not directly provide the grounded signals needed for navigation learning. This raises a central question: how can generated futures be turned into supervision for grounded trajectory prediction? We present WorldMAP, a teacher--student framework that converts world-model-generated futures into persistent semantic-spatial structure and planning-derived supervision. Its world-model-driven teacher builds semantic-spatial memory from generated videos, grounds task-relevant targets and obstacles, and produces trajectory pseudo-labels through explicit planning. A lightweight student with a multi-hypothesis trajectory head is then trained to predict navigation trajectories directly from vision-language inputs. On Target-Bench, WorldMAP achieves the best ADE and FDE among compared methods, reducing ADE by 18.0% and FDE by 42.1% relative to the best competing baseline, while lifting a small open-source VLM to DTW performance competitive with proprietary models. More broadly, the results suggest that, in embodied navigation, the value of world models may lie less in supplying action-ready imagined evidence than in synthesizing structured supervision for navigation learning.

视觉语言导航预测世界模型轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。