arXiv:2410.18647cs.RO2024-10ICLR被引 200

研究机器人操作中数据量与性能的关系,发现多样环境比数据数量更重要。

Data Scaling Laws in Imitation Learning for Robotic Manipulation

论文配图:Data Scaling Laws in Imitation Learning for Robotic Manipulation
图 1 · 摘自论文原文
  • 通过收集4万条示范数据,探索训练环境和物体多样性对泛化的影响。
  • 性能随环境与物体数量呈幂律增长,单任务策略可零样本适配同类别新物体。
  • 只需少量高效数据采集,就能实现90%以上成功率,适合实际部署场景。

数据扩展已彻底改变自然语言处理和计算机视觉领域,赋予模型卓越的泛化能力。本文探究机器人领域,特别是机器人操作任务中是否存在类似的数据扩展规律,并验证合理数据扩展能否生成单任务机器人策略,实现对同一类别中任意物体在任意环境中零样本部署。为此,我们开展了一项全面的模仿学习数据扩展实证研究。通过跨多个环境和物体收集数据,考察策略泛化性能随训练环境、物体及示范数量的变化规律。整个研究共收集超过4万条示范数据,并在严格评估协议下执行1.5万余次真实机器人滚动测试。结果揭示:策略泛化性能与环境和物体数量大致遵循幂律关系;环境与物体多样性远比示范总数重要;一旦每环境或每物体的示范数达到阈值,继续增加示范效果甚微。基于此,我们提出一种高效数据收集策略。仅用四名数据收集者工作一个下午,便收集到足够数据,使两个任务的策略在全新环境中面对未见物体时成功率接近90%。

原文摘要 · Abstract (English)

Data scaling has revolutionized fields like natural language processing and computer vision, providing models with remarkable generalization capabilities. In this paper, we investigate whether similar data scaling laws exist in robotics, particularly in robotic manipulation, and whether appropriate data scaling can yield single-task robot policies that can be deployed zero-shot for any object within the same category in any environment. To this end, we conduct a comprehensive empirical study on data scaling in imitation learning. By collecting data across numerous environments and objects, we study how a policy's generalization performance changes with the number of training environments, objects, and demonstrations. Throughout our research, we collect over 40,000 demonstrations and execute more than 15,000 real-world robot rollouts under a rigorous evaluation protocol. Our findings reveal several intriguing results: the generalization performance of the policy follows a roughly power-law relationship with the number of environments and objects. The diversity of environments and objects is far more important than the absolute number of demonstrations; once the number of demonstrations per environment or object reaches a certain threshold, additional demonstrations have minimal effect. Based on these insights, we propose an efficient data collection strategy. With four data collectors working for one afternoon, we collect sufficient data to enable the policies for two tasks to achieve approximately 90% success rates in novel environments with unseen objects.

机器人操作模仿学习数据扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。