让机器人通过自然指令精准找物放物,成功率高达75%
Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning with Dense Labeling
- 用双松弛对比学习融合正样本、未标注正样本和负样本
- 在真实场景中实现75%零样本迁移成功率
- 适合需要开放词汇理解的家用服务机器人研究
劳动力短缺推动家用服务机器人(DSR)需求增长。本文开发一种可基于开放词汇指令完成日常物品搬运任务的DSR,例如根据指令“请取挂在金属毛巾架上的右侧红色毛巾并放入左侧白色洗衣机”执行操作。该任务挑战在于从数千张室内环境图像中准确检索目标物体与容器图像,而这些图像中包含大量相似物品。为此,提出RelaX-Former模型,通过融合正样本、未标注正样本与负样本,学习多样化且鲁棒的视觉表征。在包含真实室内图像与人工标注复杂指代表达指令的数据集上评估,RelaX-Former在标准图像检索指标上优于现有基线模型。进一步在物理机器人上进行零样本迁移实验,成功实现物体到指定容器的搬运,整体成功率达75%。
原文摘要 · Abstract (English)
Growing labor shortages are increasing the demand for domestic service robots (DSRs) to assist in various settings. In this study, we develop a DSR that transports everyday objects to specified pieces of furniture based on open-vocabulary instructions. Our approach focuses on retrieving images of target objects and receptacles from pre-collected images of indoor environments. For example, given an instruction "Please get the right red towel hanging on the metal towel rack and put it in the white washing machine on the left," the DSR is expected to carry the red towel to the washing machine based on the retrieved images. This is challenging because the correct images should be retrieved from thousands of collected images, which may include many images of similar towels and appliances. To address this, we propose RelaX-Former, which learns diverse and robust representations from among positive, unlabeled positive, and negative samples. We evaluated RelaX-Former on a dataset containing real-world indoor images and human annotated instructions including complex referring expressions. The experimental results demonstrate that RelaX-Former outperformed existing baseline models across standard image retrieval metrics. Moreover, we performed physical experiments using a DSR to evaluate the performance of our approach in a zero-shot transfer setting. The experiments involved the DSR to carry objects to specific receptacles based on open-vocabulary instructions, achieving an overall success rate of 75%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。