让地图生成与导航任务一起训练,提升视觉语言导航效果。
MapDream: Task-Driven Map Learning for Vision-Language Navigation
- 用自回归方式生成鸟瞰图地图,边走边学
- 在R2R-CE和RxR-CE上达到顶尖单目表现
- 适合做视觉导航与地图联合学习的研究者
视觉-语言导航(VLN)要求智能体在部分可见的3D环境中根据自然语言指令移动。现有方法多依赖与导航策略无关的手工地图。我们提出MapDream,一种地图闭环框架,将地图构建建模为自回归鸟瞰图(BEV)图像生成。该框架联合学习地图生成与动作预测,将环境上下文压缩为仅保留导航关键属性的三通道BEV地图。通过监督预训练建立可靠的映射到控制接口,自回归设计支持通过强化学习微调实现端到端优化。在R2R-CE和RxR-CE数据集上的实验表明,该方法达到当前最优的单目性能,验证了任务驱动的生成式地图学习的有效性。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception. However, most existing approaches rely on hand-crafted maps constructed independently of the navigation policy. We argue that maps should instead be learned representations shaped directly by navigation objectives rather than exhaustive reconstructions. Based on this insight, we propose MapDream, a map-in-the-loop framework that formulates map construction as autoregressive bird's-eye-view (BEV) image synthesis. The framework jointly learns map generation and action prediction, distilling environmental context into a compact three-channel BEV map that preserves only navigation-critical affordances. Supervised pre-training bootstraps a reliable mapping-to-control interface, while the autoregressive design enables end-to-end joint optimization through reinforcement fine-tuning. Experiments on R2R-CE and RxR-CE achieve state-of-the-art monocular performance, validating task-driven generative map learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。