用图像修复思路生成自然手物交互视频,支持未知人物和物体。
iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer
- 通过两阶段扩散模型,先修复手部区域插入物体,再保证动作连贯。
- 在真实场景中表现优于现有方法,尤其在遮挡和复杂交互下更自然。
- 复用预训练模型感知能力,无需额外参数,适合长视频生成。
数字人视频生成在教育和电商领域日益流行,得益于头部-身体动画与口型同步技术的进步。然而,真实的手物交互(HOI)——即手与物体之间的复杂动态——仍面临挑战。生成自然可信的HOI重演困难,主要源于手物遮挡、物体形状与朝向差异,以及对精确物理交互的需求,更重要的是需泛化至未见过的人与物体。本文提出新框架iDiT-HOI,实现真实场景下的手物交互重演生成。具体地,我们设计统一的基于图像修复的标记处理方法Inp-TPU,结合两阶段视频扩散变换器(DiT)模型。第一阶段通过将指定物体插入手部区域生成关键帧,为后续帧提供参考;第二阶段确保手物交互的时间一致性与流畅性。本方法核心贡献在于复用预训练模型的上下文感知能力,无需引入额外参数,实现对未见物体与场景的强大泛化能力,且所提范式天然支持长视频生成。全面评估表明,该方法在复杂真实场景中显著优于现有方法,显著提升真实感与手物交互的流畅性。
原文摘要 · Abstract (English)
Digital human video generation is gaining traction in fields like education and e-commerce, driven by advancements in head-body animation and lip-syncing technologies. However, realistic Hand-Object Interaction (HOI) - the complex dynamics between human hands and objects - continues to pose challenges. Generating natural and believable HOI reenactments is difficult due to issues such as occlusion between hands and objects, variations in object shapes and orientations, and the necessity for precise physical interactions, and importantly, the ability to generalize to unseen humans and objects. This paper presents a novel framework iDiT-HOI that enables in-the-wild HOI reenactment generation. Specifically, we propose a unified inpainting-based token process method, called Inp-TPU, with a two-stage video diffusion transformer (DiT) model. The first stage generates a key frame by inserting the designated object into the hand region, providing a reference for subsequent frames. The second stage ensures temporal coherence and fluidity in hand-object interactions. The key contribution of our method is to reuse the pretrained model's context perception capabilities without introducing additional parameters, enabling strong generalization to unseen objects and scenarios, and our proposed paradigm naturally supports long video generation. Comprehensive evaluations demonstrate that our approach outperforms existing methods, particularly in challenging real-world scenes, offering enhanced realism and more seamless hand-object interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。