用单目摄像头实现全景式导航,让模型自己‘想象’周围环境。
MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
- 通过统一表征融合视觉语义与语言指令,提升动作预测可靠性。
- 在多个基准上超越现有单目方法,接近全景输入表现。
- 适合资源受限场景下的智能导航系统开发。
视觉语言导航(VLN)通常依赖全景RGB和深度信息提供丰富的空间线索,但这些传感器在实际部署中成本较高或难以获取。基于视觉语言动作(VLA)模型的近期方法虽已实现单目输入下的良好效果,但仍逊于使用全景RGB-D信息的方法。我们提出MonoDream,一种轻量级VLA框架,使单目代理能够学习统一导航表征(UNR)。该共享特征表示联合对齐导航相关的视觉语义(如全局布局、深度、未来线索)与语言驱动的动作意图,从而提升动作预测的可靠性。MonoDream进一步引入潜在全景梦境(LPD)任务,以监督UNR:仅凭单目输入,训练模型预测当前及未来步骤的全景RGB和深度观测的潜在特征。在多个VLN基准上的实验表明,MonoDream持续提升单目导航性能,并显著缩小与全景基线方法之间的差距。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results with monocular input, yet they still lag behind methods using panoramic RGB-D information. We present MonoDream, a lightweight VLA framework that enables monocular agents to learn a Unified Navigation Representation (UNR). This shared feature representation jointly aligns navigation-relevant visual semantics (e.g., global layout, depth, and future cues) and language-grounded action intent, enabling more reliable action prediction. MonoDream further introduces Latent Panoramic Dreaming (LPD) tasks to supervise the UNR, which train the model to predict latent features of panoramic RGB and depth observations at both current and future steps based on only monocular input. Experiments on multiple VLN benchmarks show that MonoDream consistently improves monocular navigation performance and significantly narrows the gap with panoramic-based agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。