arXiv:2411.01813cs.ROcs.AI2024-11CoRL被引 20

实验证明,机器人自主数据收集难以规模化,人工数据远比自动收集更有效。

So You Think You Can Scale Up Autonomous Robot Data Collection?

  • 用人类示范数据启动自主模仿学习,试图减少环境设计负担。
  • 在7个真实与仿真任务中,扩大数据量提升性能,但自动收集增益有限。
  • 揭示了自动数据收集在现实场景中的本质瓶颈,适合关注落地挑战的研究者。

机器人自主学习的长期目标是实现技能的自我获取。尽管强化学习(RL)有望实现自主数据收集,但在真实世界中仍难扩展,主要因环境设计和配置需大量人力,如重置函数或成功检测器的开发。相比之下,模仿学习(IL)几乎无需环境设计,但依赖大量人工示范。近期研究提出从少量人类示范数据出发,让自主策略自举推进。然而,本文通过一系列真实实验表明,这类方法在扩展至真实场景时,仍面临与早期强化学习相似的环境设计难题。我们对不同数据规模下多种自主IL方法进行了严格评估,涵盖7个仿真与真实任务,结果发现:虽自主收集可轻微提升性能,但单纯增加人工示范数据带来的收益远超自动收集。研究结论为负面:真实世界中规模化自主数据收集远比以往认为更困难、不切实际。这些洞察有助于未来自主学习研究聚焦核心挑战。

原文摘要 · Abstract (English)

A long-standing goal in robot learning is to develop methods for robots to acquire new skills autonomously. While reinforcement learning (RL) comes with the promise of enabling autonomous data collection, it remains challenging to scale in the real-world partly due to the significant effort required for environment design and instrumentation, including the need for designing reset functions or accurate success detectors. On the other hand, imitation learning (IL) methods require little to no environment design effort, but instead require significant human supervision in the form of collected demonstrations. To address these shortcomings, recent works in autonomous IL start with an initial seed dataset of human demonstrations that an autonomous policy can bootstrap from. While autonomous IL approaches come with the promise of addressing the challenges of autonomous RL as well as pure IL strategies, in this work, we posit that such techniques do not deliver on this promise and are still unable to scale up autonomous data collection in the real world. Through a series of real-world experiments, we demonstrate that these approaches, when scaled up to realistic settings, face much of the same scaling challenges as prior attempts in RL in terms of environment design. Further, we perform a rigorous study of autonomous IL methods across different data scales and 7 simulation and real-world tasks, and demonstrate that while autonomous data collection can modestly improve performance, simply collecting more human data often provides significantly more improvement. Our work suggests a negative result: that scaling up autonomous data collection for learning robot policies for real-world tasks is more challenging and impractical than what is suggested in prior work. We hope these insights about the core challenges of scaling up data collection help inform future efforts in autonomous learning.

机器人学习模仿学习数据收集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。