让机器人听懂模糊指令,像人一样聪明导航。
CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
- 用模仿学习融合视觉与语言指令,理解人类抽象引导。
- 在果园场景中成功率67%,远超规则系统0%表现。
- 仿真训练后真实世界部署成功率达69%,跨环境适应强。
真实场景中的机器人导航不仅需抵达目标,还需优化路径并满足特定场景需求。人类常通过口头指令或草图等抽象方式表达意图,这类指导往往缺乏细节或带有噪声。为使机器人准确理解并执行此类指令,需具备与人类一致的基本导航常识。为此,我们提出CANVAS框架,融合视觉与语言指令实现常识感知导航。其核心为模仿学习,使机器人从人类导航行为中学习。我们构建了COMMAND数据集,包含超过48小时、219公里的人类标注导航数据,用于模拟环境中训练常识导航系统。实验表明,CANVAS在所有环境下均优于强基准系统ROS NavStack,尤其在果园场景中,当后者成功率仅为0%时,CANVAS达到67%。此外,即使在未见环境中,CANVAS仍能紧密匹配人类示范与常识约束。真实世界部署结果显示,其Sim2Real迁移性能优异,总成功率达69%,验证了基于人类示范的仿真训练对现实应用的潜力。
原文摘要 · Abstract (English)
Real-life robot navigation involves more than just reaching a destination; it requires optimizing movements while addressing scenario-specific goals. An intuitive way for humans to express these goals is through abstract cues like verbal commands or rough sketches. Such human guidance may lack details or be noisy. Nonetheless, we expect robots to navigate as intended. For robots to interpret and execute these abstract instructions in line with human expectations, they must share a common understanding of basic navigation concepts with humans. To this end, we introduce CANVAS, a novel framework that combines visual and linguistic instructions for commonsense-aware navigation. Its success is driven by imitation learning, enabling the robot to learn from human navigation behavior. We present COMMAND, a comprehensive dataset with human-annotated navigation results, spanning over 48 hours and 219 km, designed to train commonsense-aware navigation systems in simulated environments. Our experiments show that CANVAS outperforms the strong rule-based system ROS NavStack across all environments, demonstrating superior performance with noisy instructions. Notably, in the orchard environment, where ROS NavStack records a 0% total success rate, CANVAS achieves a total success rate of 67%. CANVAS also closely aligns with human demonstrations and commonsense constraints, even in unseen environments. Furthermore, real-world deployment of CANVAS showcases impressive Sim2Real transfer with a total success rate of 69%, highlighting the potential of learning from human demonstrations in simulated environments for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。