用手机扫描+一段视频生成数千条机器人训练数据,无需真实机器人或物理模拟。
Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware
- 通过手机扫描和人类示范视频重建3D物体并追踪运动,生成高保真虚拟演示。
- 仅用一次人类示范训练的模型,性能媲美150次真人操作的数据训练结果。
- 适合想低成本构建机器人数据集的研究者,尤其适用于视觉-语言-动作模型训练。
机器人学习的规模化需要海量且多样化的数据集。然而,主流的数据采集方式——人工远程操控——仍成本高昂,受限于人力投入和实体机器人获取。我们提出一种新方法 Real2Render2Real(R2R2R),可在不依赖物体动力学仿真或真实机器人操控的情况下生成机器人训练数据。输入为单个或多个物体的手机扫描数据及一段人类示范视频。R2R2R 通过3D高斯点云(3DGS)重建精细的3D物体几何与外观,并追踪6自由度物体运动,生成数千条高视觉保真度的机器人无关示范。该方法将3D表示转换为网格以兼容 IsaacLab 等可扩展渲染引擎,但关闭碰撞建模。由 R2R2R 生成的机器人示范数据可直接用于基于本体感觉状态与图像观测的模型,如视觉-语言-动作模型(VLA)和模仿学习策略。物理实验表明,仅使用一次人类示范生成的数据训练的模型,其性能可达到使用150次人工遥控数据训练模型的水平。
原文摘要 · Abstract (English)
Scaling robot learning requires vast and diverse datasets. Yet the prevailing data collection paradigm-human teleoperation-remains costly and constrained by manual effort and physical robot access. We introduce Real2Render2Real (R2R2R), a novel approach for generating robot training data without relying on object dynamics simulation or teleoperation of robot hardware. The input is a smartphone-captured scan of one or more objects and a single video of a human demonstration. R2R2R renders thousands of high visual fidelity robot-agnostic demonstrations by reconstructing detailed 3D object geometry and appearance, and tracking 6-DoF object motion. R2R2R uses 3D Gaussian Splatting (3DGS) to enable flexible asset generation and trajectory synthesis for both rigid and articulated objects, converting these representations to meshes to maintain compatibility with scalable rendering engines like IsaacLab but with collision modeling off. Robot demonstration data generated by R2R2R integrates directly with models that operate on robot proprioceptive states and image observations, such as vision-language-action models (VLA) and imitation learning policies. Physical experiments suggest that models trained on R2R2R data from a single human demonstration can match the performance of models trained on 150 human teleoperation demonstrations. Project page: https://real2render2real.com
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。