用生成视频训练人形机器人,让其学会灵活抓取与全身协作操作。
RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation

- 从单视角视觉生成动作视频,通过深度重建提取关键帧。
- 在真实机器人上实现跨物体配置的泛化操作,抗干扰能力强。
- 适合想低成本获取通用操作技能的研究者与工程师。
人形机器人具备在人类环境中执行灵巧操作的潜力,但获取多样且可泛化的技能成本高昂,因硬件数据采集昂贵且标注劳动密集。近期视频生成模型为从视觉观察中合成丰富操作经验提供了契机,但将这些想象行为转化为可执行的人形机器人全身技能仍基本未被探索。本文提出RoboReact框架,可自动从单一第一人称RGB-D观测中合成全身人形机器人操作技能。RoboReact生成人类操作视频,通过深度感知3D重建提取保持几何一致性的交互关键帧,并将其重定向至高自由度人形平台,同时保留手物交互几何结构。为弥合想象规划与物理执行之间的差距,RoboReact采用在线对象中心重定位,并利用视觉-语言模型引导的精炼循环适应几何不匹配与执行偏差。经精炼的技能由全身控制器执行,实现协调的全身操作与灵巧交互。真实人形机器人实验表明,RoboReact可在多种物体配置下泛化,且在执行扰动下稳健恢复,无需遥操作或人工示范。结果凸显了生成模型、视觉-语言推理与闭环控制结合在可扩展人形技能获取中的潜力。
原文摘要 · Abstract (English)
Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive annotation. Recent advances in video generative models provide a promising opportunity to synthesize rich manipulation experiences from visual observations, but transferring such imagined behaviors into executable whole-body humanoid skills remains largely unexplored. In this work, we present RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation. RoboReact generates human manipulation videos, extracts geometry-preserving interaction keyframes through depth-aware 3D reconstruction, and retargets them to high-DoF humanoid platforms while preserving hand-object interaction geometry. To bridge the gap between imagined plans and physical execution, RoboReact performs online object-centric re-grounding and leverages a vision-language model-guided refinement loop to adapt skills under geometric mismatch and execution deviations. The refined skills are executed through a whole-body controller, enabling coordinated whole-body manipulation and dexterous interaction. Experiments on real humanoid robots demonstrate that RoboReact generalizes across diverse object configurations and robustly recovers from execution disturbances without requiring teleoperation or human demonstrations. These results highlight the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。