用众包方式收集机器人训练数据,降低人力成本。
RoboCrowd: Scaling Robot Data Collection through Crowdsourcing
- 通过奖励、趣味任务和排行榜激励公众参与机器人操作
- 两周内吸引200多人贡献超800次操作数据
- 收集的数据可提升机器人模型性能最高20%,适合大规模数据采集场景
近年来,基于大规模人类示范的模仿学习已成为训练机器人策略的有前景方法。然而,收集大量人类示范在时间和专家资源获取上存在巨大负担。本文提出RoboCrowd,一种基于众包原则与激励设计的数据收集新范式,以分担工作负荷并实现可扩展的数据采集。我们在ALOHA(Zhao等,2023)——一个支持远程操控的双臂机器人平台——基础上,探索在公共环境中众包现场示范的设计空间。提出了三类激励机制:物质奖励、内在兴趣与社会比较,分别通过实物奖励、挑战性操作任务及排行榜等实现。我们在大学咖啡厅开展为期两周的大规模实地实验,观察到显著用户参与:超过200人独立自愿提供了总计超过800个交互片段。结果验证了所提激励机制对数据数量与质量的有效影响。此外,我们证明众包数据可用于政策的预训练,在微调时相较无此数据可提升高达20%的性能。这些结果表明,通过精心设计众包与激励机制,RoboCrowd有望显著减轻机器人数据采集负担。
原文摘要 · Abstract (English)
In recent years, imitation learning from large-scale human demonstrations has emerged as a promising paradigm for training robot policies. However, the burden of collecting large quantities of human demonstrations is significant in terms of collection time and the need for access to expert operators. We introduce a new data collection paradigm, RoboCrowd, which distributes the workload by utilizing crowdsourcing principles and incentive design. RoboCrowd helps enable scalable data collection and facilitates more efficient learning of robot policies. We build RoboCrowd on top of ALOHA (Zhao et al. 2023) -- a bimanual platform that supports data collection via puppeteering -- to explore the design space for crowdsourcing in-person demonstrations in a public environment. We propose three classes of incentive mechanisms to appeal to users' varying sources of motivation for interacting with the system: material rewards, intrinsic interest, and social comparison. We instantiate these incentives through tasks that include physical rewards, engaging or challenging manipulations, as well as gamification elements such as a leaderboard. We conduct a large-scale, two-week field experiment in which the platform is situated in a university cafe. We observe significant engagement with the system -- over 200 individuals independently volunteered to provide a total of over 800 interaction episodes. Our findings validate the proposed incentives as mechanisms for shaping users' data quantity and quality. Further, we demonstrate that the crowdsourced data can serve as useful pre-training data for policies fine-tuned on expert demonstrations -- boosting performance up to 20% compared to when this data is not available. These results suggest the potential for RoboCrowd to reduce the burden of robot data collection by carefully implementing crowdsourcing and incentive design principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。