arXiv:2503.16942cs.CV2025-03CVPR被引 23

让虚拟人物自然互动真实物体,生成更逼真的手物交互视频。

Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model

  • 用分层布局表示手与物体,解耦建模提升交互精度。
  • 引入独立记忆库增强纹理细节,生成更真实的视觉效果。
  • 自适应调整布局,解决不同大小物体带来的生成偏差。

当前聚焦于口型同步和身体动作的数字人研究已难以满足产业需求,而支持与现实环境(如物体)交互的人类视频生成技术尚未充分探索。尽管手部合成已是复杂问题,但使物体与手部接触并实现自然交互更具挑战性,尤其当物体在尺寸和形状上差异显著时。为此,我们提出一种基于自适应布局引导扩散模型的视频重演框架(Re-HOLD),专注于人-物交互(HOI)。核心思路是分别采用专用于手部与物体的布局表示,实现手部建模与物体适配多样运动序列的有效解耦。为进一步提升交互生成质量,设计了针对手部与物体的交互式纹理增强模块,引入两个独立的记忆库。同时,针对跨物体重演场景,提出布局自适应调整策略,以缓解因物体尺寸差异导致的推理阶段不合理布局问题。全面的定性和定量评估表明,所提框架显著优于现有方法。

原文摘要 · Abstract (English)

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g., objects) have not been well investigated. Despite human hand synthesis already being an intricate problem, generating objects in contact with hands and their interactions presents an even more challenging task, especially when the objects exhibit obvious variations in size and shape. To tackle these issues, we present a novel video Reenactment framework focusing on Human-Object Interaction (HOI) via an adaptive Layout-instructed Diffusion model (Re-HOLD). Our key insight is to employ specialized layout representation for hands and objects, respectively. Such representations enable effective disentanglement of hand modeling and object adaptation to diverse motion sequences. To further improve the generation quality of HOI, we design an interactive textural enhancement module for both hands and objects by introducing two independent memory banks. We also propose a layout adjustment strategy for the cross-object reenactment scenario to adaptively adjust unreasonable layouts caused by diverse object sizes during inference. Comprehensive qualitative and quantitative evaluations demonstrate that our proposed framework significantly outperforms existing methods. Project page: https://fyycs.github.io/Re-HOLD.

视频生成手物交互扩散模型姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。