从一段人类操作视频中自动学习灵巧技能,实现仿真到现实的高效迁移。
Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video

- 利用基础模型重建数字孪生,提取机器人与物体运动先验。
- 通过物体关键帧优化机器人配置,提升动作对环境的影响效果。
- 结合强化学习与逆向强化学习,解决感知噪声与交互差异问题。
人类操作视频是机器人学习的便捷且直观数据源,但直接将人类灵巧性迁移到机器人仍面临感知误差与本体差异挑战。为此,我们提出 Video2Sim2Real,一个从单个真人操作视频实现自主灵巧技能获取的全栈框架。该框架首先使用现成的基础模型重建可模拟的数字孪生,并提取机器人与物体的运动先验。不同于将提取的机器人运动视为全程可靠参考,我们的核心思路是回归演示技能中最根本的监督信号:识别以物体为中心的关键帧,利用模拟器中的物体信息优化对应机器人配置,并以此作为锚点来修正机器人运动,使其最终对环境产生预期影响。为弥合剩余的仿真到现实差距,我们引入一种解耦策略,将对噪声和不完整感知的鲁棒性与手-物交互动态变化分开处理。具体而言,通过逆向强化学习从真实世界点云中重新校准机器人配置,并利用残差强化学习进行局部指端级适应,确保稳健有效的交互。最后,基于碰撞感知的运动规划模块支持对新物体构型的空间泛化。在多个日常操作任务中,Video2Sim2Real 在模拟环境中的任务成功率、安全性与轨迹连贯性均优于众多基线方法,且实现了比现有技术更优的仿真到现实迁移效果。这些结果展示了从人类视频实现自主灵巧技能获取的可行路径。
原文摘要 · Abstract (English)
Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains challenging due to perception errors and embodiment gap. To address this, we introduce Video2Sim2Real, a full-stack framework for autonomous skill acquisition from a single human manipulation video. Our framework first uses off-the-shelf foundation models to reconstruct a simulator-ready digital twin and extract robot and object motion priors. Rather than treating the extracted robot motion as a reliable reference throughout execution, our key idea is to recover and leverage the most fundamental sources of supervision from the demonstrated skill: We identify object-centric keyframes to optimize the corresponding robot configurations using object information from the simulator, and use these configurations as anchors that refine the robot motion such that it ultimately has the desired impact on the environment. To bridge the remaining sim-to-real gap, we introduce a sim-to-real strategy that decouples robustness to noisy and incomplete perception from variations in hand-object interaction dynamics. Specifically, we learn to recalibrate robot configurations from noisy real-world point clouds via IL, and leverage residual RL to perform local finger-level adaptations to ensure for robust and effective interactions. Finally, a collision-aware motion planning module enables spatial generalization to novel object configurations. Across several everyday manipulation tasks, Video2Sim2Real improves simulated task success, safety, and trajectory coherence over numerous baselines, and achieves better sim-to-real transfer than existing techniques. These results demonstrate a promising path toward autonomous dexterous skill acquisition from human videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。