arXiv:2412.01770cs.ROcs.AI2024-12被引 18

用虚拟场景+模型生成数据,让机器人学习效率随人力投入超线性提升

Robot Learning with Super-Linear Scaling

  • 通过3D重建构建虚拟场景,用强化学习加人类示范启动模拟数据收集
  • 训练中逐步用模型生成示范替代人工,实现人力消耗递减的数据积累
  • 零样本与少样本下跨场景迁移成功,视频扫描即可微调到新环境

机器人学习的规模化依赖于可高效扩展的人力驱动数据采集。本文提出一种名为CASHER(Crowdsourcing and Amortizing Human Effort for Real-to-Sim-to-Real)的模拟数据采集与学习管道,其性能随人力投入呈现超线性增长。核心思路是利用3D重建技术众包真实场景的数字孪生体,在仿真环境中大规模采集数据,而非直接在现实世界操作。初期数据采集由强化学习驱动,并以人类示范进行引导;随着通用策略在多环境下持续训练,其泛化能力被用于生成示范,逐步替代人工参与。这一过程实现了模拟环境中行为数据的持续积累,且人力投入不断减少。我们在三个真实任务上验证了CASHER的零样本与少样本缩放规律,并证明仅需视频扫描即可完成预训练策略到目标场景的微调,无需额外人工干预。

原文摘要 · Abstract (English)

Scaling robot learning requires data collection pipelines that scale favorably with human effort. In this work, we propose Crowdsourcing and Amortizing Human Effort for Real-to-Sim-to-Real(CASHER), a pipeline for scaling up data collection and learning in simulation where the performance scales superlinearly with human effort. The key idea is to crowdsource digital twins of real-world scenes using 3D reconstruction and collect large-scale data in simulation, rather than the real-world. Data collection in simulation is initially driven by RL, bootstrapped with human demonstrations. As the training of a generalist policy progresses across environments, its generalization capabilities can be used to replace human effort with model generated demonstrations. This results in a pipeline where behavioral data is collected in simulation with continually reducing human effort. We show that CASHER demonstrates zero-shot and few-shot scaling laws on three real-world tasks across diverse scenarios. We show that CASHER enables fine-tuning of pre-trained policies to a target scenario using a video scan without any additional human effort. See our project website: https://casher-robot-learning.github.io/CASHER/

机器人学习仿真训练超线性扩展数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。