生成双手操作物体的连贯视频,支持通用抓握与真实交互。
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
- 构建多层遮挡表示,提升手物交互的三维一致性。
- 融合大规模3D物体数据集,实现物体泛化抓握能力。
- 适合需要真实手物交互视频生成的研究与应用。
本文提出ManiVideo,一种从给定的手部与物体运动序列中生成一致且时间连贯的双手操作物体视频的新方法。其核心是构建多层遮挡(MLO)表示,通过无遮挡法线图和遮挡置信图学习3D遮挡关系,并以两种形式嵌入UNet结构,增强复杂手物交互的三维一致性。为实现物体的泛化抓握,引入大型3D物体数据集Objaverse,缓解视频数据稀缺问题,促进物体一致性学习。此外,提出创新训练策略,有效融合多个数据集,支持下游人本中心手物交互视频生成任务。大量实验表明,该方法不仅生成具有合理手物交互和广泛物体泛化的视频,且优于现有最先进方法。
原文摘要 · Abstract (English)
In this paper, we introduce ManiVideo, a novel method for generating consistent and temporally coherent bimanual hand-object manipulation videos from given motion sequences of hands and objects. The core idea of ManiVideo is the construction of a multi-layer occlusion (MLO) representation that learns 3D occlusion relationships from occlusion-free normal maps and occlusion confidence maps. By embedding the MLO structure into the UNet in two forms, the model enhances the 3D consistency of dexterous hand-object manipulation. To further achieve the generalizable grasping of objects, we integrate Objaverse, a large-scale 3D object dataset, to address the scarcity of video data, thereby facilitating the learning of extensive object consistency. Additionally, we propose an innovative training strategy that effectively integrates multiple datasets, supporting downstream tasks such as human-centric hand-object manipulation video generation. Through extensive experiments, we demonstrate that our approach not only achieves video generation with plausible hand-object interaction and generalizable objects, but also outperforms existing SOTA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。