用真人视频训练机器人,无需人工调奖赏,直接生成可落地的复杂动作。
SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework

- 自动从真人视频提取动作和接触信息,构建初始动作先验
- 用物理仿真优化动作,提升真实感与稳定性,支持长时序执行
- 零样本迁移至真实机器人,可自愈故障并抗干扰
构建能在真实世界中实现通用全身运动操作的人形机器人仍是重大挑战。现有方法或依赖繁琐的任务特异性奖励设计,或僵化重播参考动作无法泛化,或依赖昂贵遥操作限制扩展性。尽管真人视频蕴含丰富行为数据,但从中提取的动作先验存在遮挡、接触伪影和重定向误差,难以直接用于策略学习。为此,我们提出SUGAR——一种可扩展的数据驱动框架,将多样真人视频转化为可部署的人形运动-操作技能,无需任务特异性奖励或推理时参考动作。SUGAR分三阶段:第一,全自动流水线从非结构化视频中提取包含人-物运动轨迹和接触标签的运动学交互先验;第二,基于物理的特权精炼器使用统一模仿奖励与渐进状态池,将不完美先验转化为物理合理、高保真的技能;第三,精炼后的技能被提炼为分层自主策略,包含指令生成器与指令追踪器。我们在模拟环境和真实人形硬件上评估了六项代表性运动-操作任务。结果表明,该方法显著优于参考动作跟踪基线,性能随真人视频数据量增加而持续提升。此外,在真实世界实现了零样本迁移,具备可靠的闭环执行能力、自主故障恢复及外部扰动下的稳定长时序表现。
原文摘要 · Abstract (English)
Building humanoid robots capable of generalizable whole-body loco-manipulation in the real world remains a fundamental challenge. Existing methods either rely on laborious task-specific reward engineering, rigidly replay reference motions that fail to generalize, or depend on costly teleoperation that limits scalability. While human videos capture diverse human behaviors, motion priors inferred from them are inherently imperfect, suffering from occlusion, contact artifacts, and retargeting errors that render them unsuitable for direct policy learning. To address this, we present SUGAR, a scalable data-driven framework that converts diverse human videos into deployable humanoid loco-manipulation skills, without any task-specific reward engineering or reference-motion conditioning at inference. SUGAR proceeds in three stages. First, a fully automated pipeline extracts kinematic interaction priors including human-object motion trajectories and contact labels from unstructured human videos. Second, a privileged physics-based refiner uses a unified mimic reward and progressive state pool to transform imperfect priors into physically feasible, high-fidelity skills. Third, refined skills are distilled into a hierarchical autonomous policy consisting of a command generator and a command tracker. We evaluate SUGAR on six representative loco-manipulation tasks in simulation and real-world humanoid hardware. Our method substantially outperforms reference-tracking baselines, and performance scales clearly with the amount of human video data. It also achieves zero-shot real-world transfer with reliable closed-loop execution, autonomous failure recovery, and stable long-horizon performance under external perturbations. Project Page: https://tianshuwu.github.io/sugar-humanoid/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。