arXiv:2509.14178cs.RO2025-09被引 5

用生成视频替代真人示范,让机器人学会灵巧抓取。

\textsc{Gen2Real}: Towards Demo-Free Dexterous Manipulation by Harnessing Generated Video

  • 用视频生成+姿态估计生成手物轨迹,替代真人示范。
  • 物理一致性优化使仿真抓取成功率达77.3%。
  • 支持自然语言指令,适合零样本泛化任务。

灵巧操作仍是机器人领域的难题,主要源于收集大量人类示范成本高昂。本文提出Gen2Real,仅需一个生成视频即可驱动机器人学习技能:结合视频生成与姿态、深度估计生成手物轨迹;通过物理感知交互优化模型(PIOM)强化物理一致性;利用基于锚点的残差近端策略优化(PPO)将人类动作重定向至机械手并稳定控制。仅使用生成视频,所学策略在仿真中抓取任务成功率达77.3%,并在真实机器人上实现连贯执行。消融实验验证各组件有效性,并展示了通过自然语言直接指定任务的能力,表明Gen2Real在从想象视频到现实执行的泛化能力上具有灵活性与鲁棒性。

原文摘要 · Abstract (English)

Dexterous manipulation remains a challenging robotics problem, largely due to the difficulty of collecting extensive human demonstrations for learning. In this paper, we introduce \textsc{Gen2Real}, which replaces costly human demos with one generated video and drives robot skill from it: it combines demonstration generation that leverages video generation with pose and depth estimation to yield hand-object trajectories, trajectory optimization that uses Physics-aware Interaction Optimization Model (PIOM) to impose physics consistency, and demonstration learning that retargets human motions to a robot hand and stabilizes control with an anchor-based residual Proximal Policy Optimization (PPO) policy. Using only generated videos, the learned policy achieves a 77.3\% success rate on grasping tasks in simulation and demonstrates coherent executions on a real robot. We also conduct ablation studies to validate the contribution of each component and demonstrate the ability to directly specify tasks using natural language, highlighting the flexibility and robustness of \textsc{Gen2Real} in generalizing grasping skills from imagined videos to real-world execution.

灵巧操作视频生成零样本机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。