arXiv:2609.03591cs.RO2026-09

用1500小时人类示范数据训练出可高效扩展的双臂操作模型

Scaling Bimanual Household Manipulation from 1,500 hours of Demonstrations to On-Policy Corrections

论文配图:Scaling Bimanual Household Manipulation from 1,500 hours of Demonstrations to On-Policy Corrections
图 1 · 摘自论文原文
  • 构建高速数据管道与分阶段训练框架,提升数据利用率
  • 在真实干预数据上微调后,任务成功率随数据量持续提升
  • 开源1500小时双臂操作数据集,支持可复现研究

学习通用且鲁棒的双臂操作策略受限于高质量大规模人类示范数据的稀缺。本文发布涵盖日常家务任务的1,500小时多样化双臂操作示范数据,并基于此训练出强大的视觉-语言-动作(VLA)模型XR-2。通过专为高吞吐设计的数据流水线和精心构建的多阶段训练范式,XR-2在系统性实验中表现出色,兼具优异的性能、训练效率和高数据利用率。我们进一步探究两个关键扩展维度:专家示范数据量的变化,以及在实时人类干预生成的DAgger校正数据上进行微调。在两种设置下,任务成功率均随数据范围稳步提升,显示出当前数据规模下的清晰一致的缩放趋势。结果验证了XR-2的学习能力及所发布数据集的可扩展潜力,该数据集已开源,以支持双臂机器人操作学习的可复现研究。

原文摘要 · Abstract (English)

Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality large scale human demonstration data. In this work, we release 1,500 hours of diverse bimanual manipulation demonstrations covering everyday household tasks, and use this comprehensive corpus to train XR-2, a powerful vision-language-action (VLA) model. Enabled by a purpose built high throughput data pipeline and a carefully designed multi stage training paradigm, XR-2 attains strong manipulation performance in our systematic experiments while retaining favorable training efficiency and high data utilization. We further study two critical scaling axes: varying the amount of expert demonstration data, and post training on DAgger correction data from real time human interventions. In both settings, task success rate improves steadily over the data ranges we probe, exhibiting a clear consistent scaling trend at our current data scale. These results validate both the learning capacity of XR-2 and the promising scaling properties of the released dataset, which we open source to support reproducible research on bimanual robot manipulation learning.

双臂操作数据扩展机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。