让机器人零样本模仿生成的人类动作,实现物理合理的运动轨迹。
From Generated Human Videos to Physically Plausible Robot Trajectories
- 分两阶段处理:先将视频转为4D人体表征,再重定向到机器人形态。
- 新策略GenMimic在仿真和真实机器人上均实现稳定动作追踪。
- 构建了合成数据集GenMimicBench,用于评估零样本泛化能力。
视频生成模型在合成新颖情境下的人类动作方面快速进步,有望作为上下文感知机器人控制的高层规划器。然而,如何让类人机器人以零样本方式执行生成视频中的人类动作仍是一个关键挑战,因为生成视频通常存在噪声和形变,难以直接模仿。为此,我们提出一个两阶段流程:首先将视频像素提升为4D人体表示并重定向至类人机器人形态;其次设计GenMimic——一种基于3D关键点、具备物理感知的强化学习策略,通过对称性正则化与关键点加权追踪奖励进行训练。结果表明,GenMimic可成功模仿来自噪声生成视频的人类动作。我们构建了GenMimicBench,一个使用两种视频生成模型生成的合成人类动作数据集,涵盖多样动作与场景,建立了评估零样本泛化与策略鲁棒性的基准。大量实验显示其在仿真中优于强基线,并在不微调的情况下于Unitree G1机器人上实现连贯且物理稳定的运动追踪。本工作为实现视频生成模型作为机器人控制高层策略的潜力提供了可行路径。
原文摘要 · Abstract (English)
Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual robot control. To realize this potential, a key research question remains open: how can a humanoid execute the human actions from generated videos in a zero-shot manner? This challenge arises because generated videos are often noisy and exhibit morphological distortions that make direct imitation difficult compared to real video. To address this, we introduce a two-stage pipeline. First, we lift video pixels into a 4D human representation and then retarget to the humanoid morphology. Second, we propose GenMimic-a physics-aware reinforcement learning policy conditioned on 3D keypoints, and trained with symmetry regularization and keypoint-weighted tracking rewards. As a result, GenMimic can mimic human actions from noisy, generated videos. We curate GenMimicBench, a synthetic human-motion dataset generated using two video generation models across a spectrum of actions and contexts, establishing a benchmark for assessing zero-shot generalization and policy robustness. Extensive experiments demonstrate improvements over strong baselines in simulation and confirm coherent, physically stable motion tracking on a Unitree G1 humanoid robot without fine-tuning. This work offers a promising path to realizing the potential of video generation models as high-level policies for robot control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。