用大模型让机器人听懂自然语言指令,处理没见过的物品。
Language-Conditioned Open-Vocabulary Mobile Manipulation with Pretrained Models
- 结合大语言模型与视觉语言模型,理解自由格式指令
- 在复杂家庭场景中实现零样本泛化,多任务成功率更高
- 适合需要灵活交互的家用机器人应用
开放词汇移动操作(OVMM)在真实世界机器人应用中仍面临重大挑战,尤其涉及跨工作区处理新出现或未见过的物体时。本文提出一种新型语言条件下的开放词汇移动操作框架LOVMM,融合大型语言模型(LLM)与视觉语言模型(VLM),以应对家庭环境中多种移动操作任务。该方法能够根据自由形式的自然语言指令完成任务,例如“把办公室桌上的食品盒扔到角落的垃圾桶”或“把床边的瓶子收拾到客房的盒子里”。在复杂家庭环境中的大量仿真实验表明,LOVMM具备出色的零样本泛化能力与多任务学习能力。此外,该方法还能推广至多种桌面操作任务,并在成功率上优于其他先进方法。
原文摘要 · Abstract (English)
Open-vocabulary mobile manipulation (OVMM) that involves the handling of novel and unseen objects across different workspaces remains a significant challenge for real-world robotic applications. In this paper, we propose a novel Language-conditioned Open-Vocabulary Mobile Manipulation framework, named LOVMM, incorporating the large language model (LLM) and vision-language model (VLM) to tackle various mobile manipulation tasks in household environments. Our approach is capable of solving various OVMM tasks with free-form natural language instructions (e.g. "toss the food boxes on the office room desk to the trash bin in the corner", and "pack the bottles from the bed to the box in the guestroom"). Extensive experiments simulated in complex household environments show strong zero-shot generalization and multi-task learning abilities of LOVMM. Moreover, our approach can also generalize to multiple tabletop manipulation tasks and achieve better success rates compared to other state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。