用大模型零样本自动标注机器人长时序数据,解决语言标注稀缺问题。
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
- 基于视觉语言大模型,零样本识别场景物体与变化,分割任务并自动打标
- 在3个数据集上成功标注超115k条轨迹,覆盖430小时未标注机器人数据
- 适合需要大规模高质量语言-动作对的机器人学习研究者
开发机器人将自然语言与感知、动作关联的核心挑战在于多样机器人数据中自然语言标注的稀缺。现有机器人策略通常依赖模板化语言或昂贵的人工标注指令,难以扩展。为此,我们提出NILS:面向可扩展性的自然语言指令标注方法。NILS以零样本方式全自动标注大量未清理、长时序的机器人数据,无需人工干预。该方法结合预训练视觉-语言基础模型,实现场景中物体检测、以物体为中心的变化识别、从大量无标签交互数据中分割任务,并最终完成行为数据标注。在BridgeV2、Fractal和厨房游戏数据集上的评估表明,NILS能自主标注多样化的未标注、非结构化机器人演示数据,有效缓解众包人工标注存在的数据质量低、多样性差等问题。我们利用NILS标注了超过115,000条轨迹,来自超过430小时的机器人数据。相关自动标注代码与生成的标注数据已开源,详见:http://robottasklabeling.github.io。
原文摘要 · Abstract (English)
A central challenge towards developing robots that can relate human language to their perception and actions is the scarcity of natural language annotations in diverse robot datasets. Moreover, robot policies that follow natural language instructions are typically trained on either templated language or expensive human-labeled instructions, hindering their scalability. To this end, we introduce NILS: Natural language Instruction Labeling for Scalability. NILS automatically labels uncurated, long-horizon robot data at scale in a zero-shot manner without any human intervention. NILS combines pretrained vision-language foundation models in order to detect objects in a scene, detect object-centric changes, segment tasks from large datasets of unlabelled interaction data and ultimately label behavior datasets. Evaluations on BridgeV2, Fractal, and a kitchen play dataset show that NILS can autonomously annotate diverse robot demonstrations of unlabeled and unstructured datasets while alleviating several shortcomings of crowdsourced human annotations, such as low data quality and diversity. We use NILS to label over 115k trajectories obtained from over 430 hours of robot data. We open-source our auto-labeling code and generated annotations on our website: http://robottasklabeling.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。