用生成视频让机器人零样本完成复杂交互,避免动作错位问题。
Morphology-Consistent Humanoid Interaction through Robot-Centric Video Synthesis
- 以机器人自身坐标系生成动作视频,直接提取可行关节轨迹。
- 在4项任务中成功率达37.5%,传统方法为0%。
- 适合无额外训练的移动人形机器人快速部署交互能力。
赋予人形机器人多样交互能力通常需要大量策略训练或显式的人体到机器人的动作迁移。然而,基于学习的策略面临高昂的数据采集成本;而迁移方法依赖人体中心的姿态估计(如SMPL),引入形态差异。骨骼尺度不匹配导致映射至机器人时产生严重空间错位,影响交互成功率。本文提出Dream2Act,一种以机器人为中心的框架,通过生成视频实现零样本交互。给定机器人第三人称图像与目标物体,该框架利用视频生成模型构想机器人完成任务的形态一致动作。我们采用高保真姿态提取系统,从合成动作中恢复物理可行的机器人本体关节轨迹,并通过通用全身控制器执行。整个过程严格在机器人本体坐标系内运行,避免迁移误差且无需任务特异性策略训练。我们在Unitree G1上评估了四个全身移动交互任务:踢球、沙发坐姿、包击打和抱箱。Dream2Act整体成功率37.5%,而传统迁移方法为0%。由于形态差距导致物理接触错误(尤其在运动中加剧),而Dream2Act保持机器人一致的空间对齐,实现了可靠接触形成和显著更高的任务完成率。
原文摘要 · Abstract (English)
Equipping humanoid robots with versatile interaction skills typically requires either extensive policy training or explicit human-to-robot motion retargeting. However, learning-based policies face prohibitive data collection costs. Meanwhile, retargeting relies on human-centric pose estimation (e.g., SMPL), introducing a morphology gap. Skeletal scale mismatches result in severe spatial misalignments when mapped to robots, compromising interaction success. In this work, we propose Dream2Act, a robot-centric framework enabling zero-shot interaction through generative video synthesis. Given a third-person image of the robot and target object, our framework leverages video generation models to envision the robot completing the task with morphology-consistent motion. We employ a high-fidelity pose extraction system to recover physically feasible, robot-native joint trajectories from these synthesized dreams, subsequently executed via a general-purpose whole-body controller. Operating strictly within the robot-native coordinate space, Dream2Act avoids retargeting errors and eliminates task-specific policy training. We evaluate Dream2Act on the Unitree G1 across four whole-body mobile interaction tasks: ball kicking, sofa sitting, bag punching, and box hugging. Dream2Act achieves a 37.5% overall success rate, compared to 0% for conventional retargeting. While retargeting fails to establish correct physical contacts due to the morphology gap (with errors compounded during locomotion), Dream2Act maintains robot-consistent spatial alignment, enabling reliable contact formation and substantially higher task completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。