同步扩散生成手物交互视频与动作,兼顾视觉真实与物理合理。
SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
- 用同步扩散融合视觉、运动与语义信息,实现端到端生成。
- 在未见过的真实场景中生成高保真、动态合理的手物交互序列。
- 无需预设物体模型或姿态引导,适合复杂交互场景建模。
手物交互(HOI)生成具有广泛应用潜力。然而,现有3D HOI动作生成方法严重依赖预定义3D物体模型和实验室采集的动作数据,限制了泛化能力;而HOI视频生成方法虽注重像素级视觉保真度,常牺牲物理合理性。我们提出一种新框架,通过在同步扩散过程中融合视觉先验与动态约束,同时生成高质量的HOI视频与动作。为对齐异构的语义、外观与运动特征,方法采用三模态自适应调制与3D全注意力机制,建模模态间及模态内依赖关系。进一步引入视觉感知的3D交互扩散模型,从同步扩散输出直接生成显式3D交互序列,并反馈形成闭环循环。该架构摆脱对预设物体模型或显式姿态引导的依赖,显著提升视频与动作的一致性。实验表明,本方法在生成高保真、动态合理的HOI序列方面优于当前最优方法,且在未见真实场景中具备出色泛化能力。
原文摘要 · Abstract (English)
Hand-Object Interaction (HOI) generation has significant application potential. However, current 3D HOI motion generation approaches heavily rely on predefined 3D object models and lab-captured motion data, limiting generalization capabilities. Meanwhile, HOI video generation methods prioritize pixel-level visual fidelity, often sacrificing physical plausibility. Recognizing that visual appearance and motion patterns share fundamental physical laws in the real world, we propose a novel framework that combines visual priors and dynamic constraints within a synchronized diffusion process to generate the HOI video and motion simultaneously. To integrate the heterogeneous semantics, appearance, and motion features, our method implements tri-modal adaptive modulation for feature aligning, coupled with 3D full-attention for modeling inter- and intra-modal dependencies. Furthermore, we introduce a vision-aware 3D interaction diffusion model that generates explicit 3D interaction sequences directly from the synchronized diffusion outputs, then feeds them back to establish a closed-loop feedback cycle. This architecture eliminates dependencies on predefined object models or explicit pose guidance while significantly enhancing video-motion consistency. Experimental results demonstrate our method's superiority over state-of-the-art approaches in generating high-fidelity, dynamically plausible HOI sequences, with notable generalization capabilities in unseen real-world scenarios. Project page at https://github.com/Droliven/SViMo_project.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。