让机器人在有移动人类的室内环境中听懂指令导航
AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
- 引入动态人类障碍,模拟真实室内导航场景
- 构建支持动画人类的仿真环境与新数据集
- 提供时间冻结机制,确保实验可复现
视觉语言导航要求机器人根据自然语言指令在真实环境中导航。以往研究多集中于静态场景,但现实导航常需应对动态人类障碍。为此,我们提出自适应视觉语言导航(AdaVLN),要求机器人在充满动态移动人类的复杂3D室内环境中导航,显著提升任务真实性。为支持该任务,我们构建了AdaVLN仿真器和AdaR2R数据集。仿真器可将全动画人类模型轻松融入Matterport3D等常见数据集。我们还引入“时间冻结”机制,在代理推理期间暂停世界状态更新,确保不同硬件上的公平比较与实验复现性。我们在该任务上评估多个基线模型,分析由AdaVLN带来的独特挑战,并验证其缩小视觉语言导航领域“仿真到现实”差距的潜力。
原文摘要 · Abstract (English)
Visual Language Navigation is a task that challenges robots to navigate in realistic environments based on natural language instructions. While previous research has largely focused on static settings, real-world navigation must often contend with dynamic human obstacles. Hence, we propose an extension to the task, termed Adaptive Visual Language Navigation (AdaVLN), which seeks to narrow this gap. AdaVLN requires robots to navigate complex 3D indoor environments populated with dynamically moving human obstacles, adding a layer of complexity to navigation tasks that mimic the real-world. To support exploration of this task, we also present AdaVLN simulator and AdaR2R datasets. The AdaVLN simulator enables easy inclusion of fully animated human models directly into common datasets like Matterport3D. We also introduce a "freeze-time" mechanism for both the navigation task and simulator, which pauses world state updates during agent inference, enabling fair comparisons and experimental reproducibility across different hardware. We evaluate several baseline models on this task, analyze the unique challenges introduced by AdaVLN, and demonstrate its potential to bridge the sim-to-real gap in VLN research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。