数据多样性不等于更好,任务多样性才是机器人学习的关键。
Is Diversity All You Need for Scalable Robotic Manipulation?
- 通过三维度分析数据多样性,发现任务多样性比示范数量更重要。
- 单机器人高质量数据预训练可高效迁移至多平台,优于多机器人数据。
- 提出去偏方法缓解速度歧义,性能提升15%,相当于2.5倍数据量。
数据规模推动了自然语言处理和计算机视觉领域基础模型的成功,但机器人操作中的有效数据扩展原则仍不明确。本文通过考察任务、机器人本体和专家三维度的数据多样性,挑战了‘越多样越好’的传统观念。在多种机器人平台的实验中发现:(1)任务多样性比每任务示范数量更重要,能更好促进从预训练到新任务的迁移;(2)多本体数据预训练对跨平台迁移非必需——仅用高质量单本体数据训练的模型,在微调阶段展现出更优的可扩展性;(3)专家多样性(如操作偏好与演示随机性)会干扰策略学习,速度多模态是关键影响因素。基于此,提出分布去偏方法以缓解速度歧义,所得GO-1-Pro实现15%性能提升,等效于使用2.5倍预训练数据。研究为高效扩展机器人操作数据集提供了新视角与实践指导。
原文摘要 · Abstract (English)
Data scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles of effective data scaling in robotic manipulation remain insufficiently understood. In this work, we investigate the nuanced role of data diversity in robot learning by examining three critical dimensions-task (what to do), embodiment (which robot to use), and expert (who demonstrates)-challenging the conventional intuition of "more diverse is better". Throughout extensive experiments on various robot platforms, we reveal that (1) task diversity proves more critical than per-task demonstration quantity, benefiting transfer from diverse pre-training tasks to novel downstream scenarios; (2) multi-embodiment pre-training data is optional for cross-embodiment transfer-models trained on high-quality single-embodiment data can efficiently transfer to different platforms, showing more desirable scaling property during fine-tuning than multi-embodiment pre-trained models; and (3) expert diversity, arising from individual operational preferences and stochastic variations in human demonstrations, can be confounding to policy learning, with velocity multimodality emerging as a key contributing factor. Based on this insight, we propose a distribution debiasing method to mitigate velocity ambiguity, the yielding GO-1-Pro achieves substantial performance gains of 15%, equivalent to using 2.5 times pre-training data. Collectively, these findings provide new perspectives and offer practical guidance on how to scale robotic manipulation datasets effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。