开源13万+任务数据集,助力机器人行为克隆研究
Scalable Behavior Cloning with Open Data, Training, and Evaluation

- 构建130K条真实操作数据集,支持高效行为克隆训练
- 400小时仿真-真实协同数据,实现低成本模型验证
- 开放全套工具链,适合机器人、强化学习研究者使用
我们提出ABC,一个完整的开源操作技能行为克隆系统。核心是目前最大的开源遥操作数据集ABC-130K,包含超过130,000个任务实例,覆盖195种多样化任务,总计3,500小时真实操作数据。同时公开硬件配置、训练基础设施与仿真流程。我们还发布400小时仿真-遥操作数据,并提供联合训练方案,使仿真与真实世界评估结果高度相关,可在不进行昂贵真实测试前有效评估模型设计与训练策略。通过对比Diffusion Transformers(DiT)与视觉-语言-动作(VLA)模型的多种架构与训练方法,验证其在真实场景中的表现,成功完成盒折叠、从钱包中取出信用卡等灵巧操作。该工具包为研究者提供可复现的基线,推动行为克隆社区协同发展。
原文摘要 · Abstract (English)
We introduce ABC, a fully open-source stack for manipulation with behavior cloning. At its core is ABC-130K: the largest open-source teleoperation dataset to date, featuring 3,500 hours of data spanning over 130K episodes across 195 diverse tasks. Furthermore, we open-source our accessible hardware setup, training infrastructure, and simulation pipeline. We also release 400 hours of sim-teleop data and provide a co-training recipe that produces correlated simulation and real-world evaluation, offering a reliable proxy for ablating model-design and training decisions before costly real-world evaluation. We explore various training recipes and compare common architectural choices for Diffusion Transformers (DiT) and Vision-Language-Action (VLA) models, grounding our findings in real-world evaluations. The resulting policies successfully execute dexterous tasks such as box folding and extracting credit cards from wallets. By providing a reproducible toolkit, we aim to place researchers on an equal footing, establishing the necessary foundation to learn the ABCs of Behavior Cloning together as a community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。