用3D高斯溅射统一生成视觉与语义信息,提升连续环境导航能力
UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation
- 基于3DGS联合渲染全景图像与语义特征
- 在R2R、REVERIE等基准上超越现有方法
- 适合需要强泛化能力的视觉语言导航研究
视觉-语言导航(VLN)中,智能体需根据指令到达目标位置。与离散环境中的路径导航不同,连续环境下的视觉-语言导航(VLN-CE)更具挑战性,因智能体可自由移动至无遮挡区域,但易受视觉遮挡或盲区影响。近期方法尝试通过预测未来环境的视觉图像或语义特征来应对,而非仅依赖当前观测。然而,这些基于RGB或特征的方法缺乏直观的外观信息或高层次语义复杂性,难以有效支持导航。为此,本文提出一种新的通用3D高斯溅射预训练范式UnitedVLN,通过统一渲染高保真360°视觉图像和语义特征,使智能体能更有效地探索未来环境。UnitedVLN采用“搜索-查询采样”与“分离-统一渲染”两种关键策略,高效利用神经基元,融合外观与语义信息,实现更鲁棒的导航。大量实验表明,UnitedVLN在现有VLN-CE基准(如R2R、REVERIE)上优于当前最先进方法。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN), where an agent follows instructions to reach a target destination, has recently seen significant advancements. In contrast to navigation in discrete environments with predefined trajectories, VLN in Continuous Environments (VLN-CE) presents greater challenges, as the agent is free to navigate any unobstructed location and is more vulnerable to visual occlusions or blind spots. Recent approaches have attempted to address this by imagining future environments, either through predicted future visual images or semantic features, rather than relying solely on current observations. However, these RGB-based and feature-based methods lack intuitive appearance-level information or high-level semantic complexity crucial for effective navigation. To overcome these limitations, we introduce a novel, generalizable 3DGS-based pre-training paradigm, called UnitedVLN, which enables agents to better explore future environments by unitedly rendering high-fidelity 360 visual images and semantic features. UnitedVLN employs two key schemes: search-then-query sampling and separate-then-united rendering, which facilitate efficient exploitation of neural primitives, helping to integrate both appearance and semantic information for more robust navigation. Extensive experiments demonstrate that UnitedVLN outperforms state-of-the-art methods on existing VLN-CE benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。