让低视角机器人更好理解人类指令,提升真实环境导航能力。
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
- 用加权历史视觉信息增强时空上下文,缓解视角差异导致的特征混淆。
- 在模拟与真实环境中均显著提升导航准确率,尤其在四足机器人上表现优异。
- 首次揭示视觉高度不匹配对导航泛化的影响,适合机器人部署研究者参考。
视觉-语言导航(VLN)使智能体能将时序视觉观测与对应指令关联,实现序列决策。然而,泛化仍是持续挑战,尤其在视觉场景多样或从仿真环境迁移到真实世界时。本文针对人类指令与四足机器人低视角观测之间的不匹配问题,提出地面视角导航(GVNav)方法,首次揭示了视觉观测高度变化对VLN泛化能力的影响。该方法利用加权的历史视觉观测构建丰富时空上下文,通过为不同视角下的相同特征分配合适权重,有效缓解单元内特征冲突,帮助低视角机器人克服视觉遮挡和感知偏差。此外,我们从HM3D与Gibson数据集迁移连接图作为额外空间先验,增强对现实场景的表征能力,显著提升航点预测器在真实环境中的性能与泛化性。大量实验表明,所提GVNav方法在模拟环境及四足机器人真实部署中均有显著提升。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, generalization remains a persistent challenge, particularly when dealing with visually diverse scenes or transitioning from simulated environments to real-world deployment. In this paper, we address the mismatch between human-centric instructions and quadruped robots with a low-height field of view, proposing a Ground-level Viewpoint Navigation (GVNav) approach to mitigate this issue. This work represents the first attempt to highlight the generalization gap in VLN across varying heights of visual observation in realistic robot deployments. Our approach leverages weighted historical observations as enriched spatiotemporal contexts for instruction following, effectively managing feature collisions within cells by assigning appropriate weights to identical features across different viewpoints. This enables low-height robots to overcome challenges such as visual obstructions and perceptual mismatches. Additionally, we transfer the connectivity graph from the HM3D and Gibson datasets as an extra resource to enhance spatial priors and a more comprehensive representation of real-world scenarios, leading to improved performance and generalizability of the waypoint predictor in real-world environments. Extensive experiments demonstrate that our Ground-level Viewpoint Navigation (GVnav) approach significantly improves performance in both simulated environments and real-world deployments with quadruped robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。