提出真实机器人场景下的视觉语言导航评估平台,揭示现有模型在物理部署中的严重退化。
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
- 构建支持多种机器人的物理仿真平台VLN-PE,实现真实运动挑战下的评估。
- 发现模型在光照变化和视野受限下性能显著下降,腿式机器人易碰撞摔倒。
- 提供可扩展框架,适合研究跨体感适应与鲁棒导航的学者使用。
当前视觉语言导航(VLN)进展虽令人鼓舞,但其理想化假设忽略了机器人运动与控制的物理现实挑战。为此,我们提出VLN-PE,一个支持人形、四足和轮式机器人的物理真实型VLN平台。首次系统评估了多种以自我为中心的VLN方法在不同技术路径下的表现:包括单步离散动作预测的分类模型、密集路径点预测的扩散模型,以及无需训练、结合地图与路径规划的大型语言模型(LLM)。结果表明,受限的观测视野、环境光照变化及物理挑战(如碰撞与跌倒)导致性能大幅下降,且腿式机器人在复杂环境中存在明显运动限制。VLN-PE高度可扩展,支持新增场景(如超越MP3D),推动更全面的VLN评估。尽管当前模型在物理部署中泛化能力弱,但该平台为提升跨体感适应性提供了新路径。代码已公开。
原文摘要 · Abstract (English)
Recent Vision-and-Language Navigation (VLN) advancements are promising, but their idealized assumptions about robot movement and control fail to reflect physically embodied deployment challenges. To bridge this gap, we introduce VLN-PE, a physically realistic VLN platform supporting humanoid, quadruped, and wheeled robots. For the first time, we systematically evaluate several ego-centric VLN methods in physical robotic settings across different technical pipelines, including classification models for single-step discrete action prediction, a diffusion model for dense waypoint prediction, and a train-free, map-based large language model (LLM) integrated with path planning. Our results reveal significant performance degradation due to limited robot observation space, environmental lighting variations, and physical challenges like collisions and falls. This also exposes locomotion constraints for legged robots in complex environments. VLN-PE is highly extensible, allowing seamless integration of new scenes beyond MP3D, thereby enabling more comprehensive VLN evaluation. Despite the weak generalization of current models in physical deployment, VLN-PE provides a new pathway for improving cross-embodiment's overall adaptability. We hope our findings and tools inspire the community to rethink VLN limitations and advance robust, practical VLN models. The code is available at https://crystalsixone.github.io/vln_pe.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。