arXiv:2411.07848cs.ROcs.CV2024-11被引 5

让机器人零样本理解并执行以物体为中心的自然语言指令。

Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation

  • 用语言推理构建因子图,实现指令与地图的零样本关联。
  • 在新环境中边建图边导航,真实世界表现稳定可靠。
  • 适合需快速部署的机器人导航场景,尤其擅长物体导向任务。

大型场景如多层住宅可通过联合估计地标与机器人位姿的3D图结构进行鲁棒高效建图,该技术广泛应用于无人机和扫地机器人等商用设备。本文提出语言推理因子图(LIFGIF),一种零样本方法,用于将自然语言指令定位到此类地图中。LIFGIF还包含一个在新环境构建地图的同时执行自然语言导航指令的策略,实现了物理世界中的稳健导航性能。为评估LIFGIF,我们构建了新的数据集——物体中心视觉语言导航(OC-VLN),用于评估以物体为中心的自然语言导航指令的定位能力。与两个相关任务的先进零样本基线(物体目标导航和视觉语言导航)相比,LIFGIF在所有评测指标上均表现更优。最后,我们在波士顿动力Spot机器人上成功验证了LIFGIF在真实世界中实现零样本物体中心指令跟随的有效性。

原文摘要 · Abstract (English)

Large scale scenes such as multifloor homes can be robustly and efficiently mapped with a 3D graph of landmarks estimated jointly with robot poses in a factor graph, a technique commonly used in commercial robots such as drones and robot vacuums. In this work, we propose Language-Inferred Factor Graph for Instruction Following (LIFGIF), a zero-shot method to ground natural language instructions in such a map. LIFGIF also includes a policy for following natural language navigation instructions in a novel environment while the map is constructed, enabling robust navigation performance in the physical world. To evaluate LIFGIF, we present a new dataset, Object-Centric VLN (OC-VLN), in order to evaluate grounding of object-centric natural language navigation instructions. We compare to two state-of-the-art zero-shot baselines from related tasks, Object Goal Navigation and Vision Language Navigation, to demonstrate that LIFGIF outperforms them across all our evaluation metrics on OCVLN. Finally, we successfully demonstrate the effectiveness of LIFGIF for performing zero-shot object-centric instruction following in the real world on a Boston Dynamics Spot robot.

机器人导航零样本语言理解物体导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。