用网页视频训练视觉语言导航,让模型学会从真实场景中理解空间。
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
- 从网络房间视频中提取动作与描述轨迹,构建大规模真实场景数据集。
- 引入隐式几何表示,直接从图像中获取空间信息,避免3D重建失败。
- 在多个基准上达到新最好效果,支持零样本导航,适合真实世界应用。
视觉-语言导航(VLN)长期受限于模拟器生成数据集的多样性与可扩展性不足,难以反映真实环境复杂性。为此,我们提出一个基于网络房间游览视频的大规模视频指令框架,使智能体能从自然的人类行走示范中学习,覆盖多样且真实的室内场景。与现有数据集不同,该框架整合了丰富语义的开放描述轨迹和重建的三维动作轨迹,提供更全面的空间与语义监督。本工作的关键创新在于引入隐式几何表示,直接从RGB帧中提取空间线索,无需依赖易出错的三维重建。该方法显著提升数据利用率,缓解重建失败问题,并释放大量以往无法使用的视频数据。在多个VLN基准(CVDN、SOON、R2R、REVERIE)上的实验表明,该方法不仅达到新的最佳性能,还实现了鲁棒的零样本导航能力。通过连接大规模网络视频与隐式空间推理,本工作推动具身导航向更可扩展、泛化更强和真实可用的方向发展。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) has long been constrained by the limited diversity and scalability of simulator-curated datasets, which fail to capture the complexity of real-world environments. To overcome this limitation, we introduce a large-scale video-instruction framework derived from web-based room tour videos, enabling agents to learn from natural human walking demonstrations in diverse, realistic indoor settings. Unlike existing datasets, our framework integrates both open-ended description-enriched trajectories and action-enriched trajectories reconstructed in 3D, providing richer spatial and semantic supervision. A key extension in this work is the incorporation of implicit geometry representations, which extract spatial cues directly from RGB frames without requiring fragile 3D reconstruction. This approach substantially improves data utilization, alleviates reconstruction failures, and unlocks large portions of previously unusable video data. Comprehensive experiments across multiple VLN benchmarks (CVDN, SOON, R2R, and REVERIE) demonstrate that our method not only sets new state-of-the-art performance but also enables the development of robust zero-shot navigation agents. By bridging large-scale web videos with implicit spatial reasoning, this work advances embodied navigation towards more scalable, generalizable, and real-world applicable solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。