用人类示范视频生成机器人可用的训练数据,提升泛化能力
MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
- 将人类操作视频通过视觉、视角和动作对齐,转为机器人可执行的指令
- 在6个任务上,仅用合成数据训练的模型成功率比纯机器人数据高14.7%
- 适合希望低成本扩展机器人训练数据的研究者与开发者
视觉语言动作(VLA)模型的泛化能力依赖于多样化的训练数据,但获取具身机器人交互数据成本高昂。相比之下,人类示范视频更易获取且成本低,近期研究也证实其在训练VLA模型中的有效性。然而,人类视频与机器人执行视频之间仍存在显著领域差异,包括不稳定的摄像头视角、人手与机械臂的视觉差异以及运动动态不同。为此,我们提出MimicDreamer框架,通过联合对齐视觉、视角和动作,将快速、低成本的人类示范转化为机器人可用的监督信号,直接支持策略训练。视觉对齐方面,提出H2R Aligner,一种视频扩散模型,通过迁移人类操作视频中的动作,生成高保真机器人示范视频;视角稳定方面,提出EgoStabilizer,通过单应性变换标准化第一人称视频,并使用图像修复填补形变和遮挡造成的缺失;动作对齐方面,将人类手部轨迹映射至机器人坐标系,采用约束逆运动学求解器生成低抖动、高精度姿态跟踪的可行关节命令。实证表明,仅使用合成人类-机器人视频训练的VLA模型即可在真实机器人上实现少样本执行。此外,利用人类数据进行规模扩展显著优于仅使用真实机器人数据的模型;本方法在六个代表性操作任务上平均成功率提升14.7%。
原文摘要 · Abstract (English)
Vision Language Action (VLA) models derive their generalization capability from diverse training data, yet collecting embodied robot interaction data remains prohibitively expensive. In contrast, human demonstration videos are far more scalable and cost-efficient to collect, and recent studies confirm their effectiveness in training VLA models. However, a significant domain gap persists between human videos and robot-executed videos, including unstable camera viewpoints, visual discrepancies between human hands and robotic arms, and differences in motion dynamics. To bridge this gap, we propose MimicDreamer, a framework that turns fast, low-cost human demonstrations into robot-usable supervision by jointly aligning vision, viewpoint, and actions to directly support policy training. For visual alignment, we propose H2R Aligner, a video diffusion model that generates high-fidelity robot demonstration videos by transferring motion from human manipulation footage. For viewpoint stabilization, EgoStabilizer is proposed, which canonicalizes egocentric videos via homography and inpaints occlusions and distortions caused by warping. For action alignment, we map human hand trajectories to the robot frame and apply a constrained inverse kinematics solver to produce feasible, low-jitter joint commands with accurate pose tracking. Empirically, VLA models trained purely on our synthesized human-to-robot videos achieve few-shot execution on real robots. Moreover, scaling training with human data significantly boosts performance compared to models trained solely on real robot data; our approach improves the average success rate by 14.7\% across six representative manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。