用合成视频教会机器人精准抓取陌生物体,无需真实动作数据。
DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation

- 通过2D视频+3D人体追踪融合,提升生成视频的物理精度。
- 在陌生物体上实现零样本泛化,手部操作成功率显著高于基线方法。
- 适合需要多样动作指令的复杂交互任务,如多物协作与文本控制。
近期视频生成模型可合成涵盖多种场景与物体类别的逼真人-物交互视频,包括难以通过动捕系统捕捉的精细操作。尽管这些合成视频蕴含丰富的交互知识,可用于精细机器人操作的运动规划,但其物理保真度有限且仅为二维,难以直接用于基于物理的角色控制模仿。本文提出DeVI(Dexterous Video Imitation)框架,利用文本条件合成视频,实现对未见目标物体的物理合理精细操作控制。为克服生成2D线索的不精确性,引入结合3D人体追踪与鲁棒2D物体追踪的混合跟踪奖励。与依赖高质量3D运动示范的方法不同,DeVI仅需生成视频即可实现跨物体、跨交互类型的零样本泛化。大量实验表明,DeVI在建模精细手-物交互方面优于现有基于3D示范的模仿方法。进一步验证了其在多物体场景和文本驱动动作多样性中的有效性,展示了视频作为人-物交互感知运动规划器的优势。
原文摘要 · Abstract (English)
Recent advances in video generative models enable the synthesis of realistic human-object interaction videos across a wide range of scenarios and object categories, including complex dexterous manipulations that are difficult to capture with motion capture systems. While the rich interaction knowledge embedded in these synthetic videos holds strong potential for motion planning in dexterous robotic manipulation, their limited physical fidelity and purely 2D nature make them difficult to use directly as imitation targets in physics-based character control. We present DeVI (Dexterous Video Imitation), a novel framework that leverages text-conditioned synthetic videos to enable physically plausible dexterous agent control for interacting with unseen target objects. To overcome the imprecision of generative 2D cues, we introduce a hybrid tracking reward that integrates 3D human tracking with robust 2D object tracking. Unlike methods relying on high-quality 3D kinematic demonstrations, DeVI requires only the generated video, enabling zero-shot generalization across diverse objects and interaction types. Extensive experiments demonstrate that DeVI outperforms existing approaches that imitate 3D human-object interaction demonstrations, particularly in modeling dexterous hand-object interactions. We further validate the effectiveness of DeVI in multi-object scenes and text-driven action diversity, showcasing the advantage of using video as an HOI-aware motion planner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。