让导航模型理解动态场景的几何结构,提升真实环境下的泛化能力。
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
- 通过跨分支融合动态几何基础模型,显式建模3D空间关系。
- 在多个基准上达到当前最优性能,真实环境鲁棒性强。
- 自适应剪枝策略降低长时程导航推理开销,无需位姿信息。
视觉语言导航(VLN)要求智能体结合视觉观测与语言指令,在未见环境中完成导航任务。现有方法多依赖静态场景假设,在动态真实场景中泛化能力不足。为此,本文提出DyGeoVLN,一种动态几何感知的VLN框架。通过跨分支特征融合,将动态几何基础模型注入原有框架,实现显式的三维空间表征与视觉-语义联合推理。为高效压缩长时程动态导航中的历史记忆,进一步设计了一种无位姿、自适应分辨率的令牌剪枝策略,可剔除时空冗余信息,显著降低推理成本。大量实验表明,该方法在多个基准上取得当前最优表现,且在真实场景中展现出优异鲁棒性。
原文摘要 · Abstract (English)
Vision-language Navigation (VLN) requires an agent to understand visual observations and language instructions to navigate in unseen environments. Most existing approaches rely on static scene assumptions and struggle to generalize in dynamic, real-world scenarios. To address this challenge, we propose DyGeoVLN, a dynamic geometry-aware VLN framework. Our method infuses a dynamic geometry foundation model into the VLN framework through cross-branch feature fusion to enable explicit 3D spatial representation and visual-semantic reasoning. To efficiently compress historical token information in long-horizon, dynamic navigation, we further introduce a novel pose-free and adaptive-resolution token-pruning strategy. This strategy can remove spatio-temporal redundant tokens to reduce inference cost. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on multiple benchmarks and exhibits strong robustness in real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。