arXiv:2409.15172cs.ROcs.AI2024-09ICRA被引 4

用互联网数据选机器人烹饪动作模板,成功率79%

Skills Made to Order: Efficient Acquisition of Robot Cooking Skills Guided by Multiple Forms of Internet Data

  • 用大模型和视频特征从网络数据中挑选合适动作模板
  • 光学流编码比传统视频编码效果好,仅用少量数据超越强基线
  • 融合多源网络数据可显著提升机器人技能学习效率

本研究探索利用多种互联网数据源,从一组预设机器人行为模板中选择适合执行接触密集型工具使用技能的方法。以往基于互联网数据学习高接触性技能面临物理信息缺失的问题,如接触存在、位置、区域与力的大小。此前工作通常利用互联网数据和训练于其上的基础模型生成低层机器人行为。我们提出假设:这些数据和模型更适合在一系列基础行为中进行选择。本文探索三种模板选择方法:调用大语言模型、使用预训练视频编码器比较机器人执行视频与检索到的人类视频,以及使用在互联网数据上训练的光流编码器进行相同比较。结果表明,尽管缺乏视觉信息,大语言模型在模板选择上表现惊人;光流编码显著优于使用千倍更多数据训练的视频编码器;且不同形式互联网数据间存在重要协同效应。通过挖掘这些协同作用,我们构建了基于多源互联网数据的模板选择器,在16种涉及工具使用的烹饪技能上达到79%的成功率。

原文摘要 · Abstract (English)

This study explores the utility of various internet data sources to select among a set of template robot behaviors to perform skills. Learning contact-rich skills involving tool use from internet data sources has typically been challenging due to the lack of physical information such as contact existence, location, areas, and force in this data. Prior works have generally used internet data and foundation models trained on this data to generate low-level robot behavior. We hypothesize that these data and models may be better suited to selecting among a set of basic robot behaviors to perform these contact-rich skills. We explore three methods of template selection: querying large language models, comparing video of robot execution to retrieved human video using features from a pretrained video encoder common in prior work, and performing the same comparison using features from an optic flow encoder trained on internet data. Our results show that LLMs are surprisingly capable template selectors despite their lack of visual information, optical flow encoding significantly outperforms video encoders trained with an order of magnitude more data, and important synergies exist between various forms of internet data for template selection. By exploiting these synergies, we create a template selector using multiple forms of internet data that achieves a 79\% success rate on a set of 16 different cooking skills involving tool-use.

机器人技能多模态学习大模型应用技能迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。