arXiv:2511.19319cs.CV2025-11被引 1

同步生成多视角手物交互视频与4D动态,解决真实场景下动作不自然问题。

SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis

  • 联合扩散模型同步生成多视角视频与中间运动轨迹。
  • 通过闭环反馈实现2D外观与4D动态的精准对齐,提升一致性。
  • 无需实验室3D数据,适用于真实世界手物交互生成。

手-物交互(HOI)生成在动画与机器人领域至关重要。现有视频方法多为单视角,难以感知完整三维几何,常导致形变或不合理运动;而3D方法依赖实验室采集的高质量3D数据,泛化性差。为此,我们提出SyncMV4D,首个联合生成同步多视角HOI视频与4D运动的模型,统一视觉先验、运动动力学与多视角几何。框架包含两大创新:(1) 多视角联合扩散(MJD)模型,共同生成HOI视频与中间运动;(2) 扩散点对齐器(DPA),将粗略中间运动优化为全局一致的4D度量点轨迹。通过闭环互增强机制,生成视频指导4D运动精修,对齐点轨迹又反向引导下一步联合生成。实验表明,该方法在视觉真实感、运动合理性与多视角一致性上均优于现有最优方案。

原文摘要 · Abstract (English)

Hand-Object Interaction (HOI) generation plays a critical role in advancing applications across animation and robotics. Current video-based methods are predominantly single-view, which impedes comprehensive 3D geometry perception and often results in geometric distortions or unrealistic motion patterns. While 3D HOI approaches can generate dynamically plausible motions, their dependence on high-quality 3D data captured in controlled laboratory settings severely limits their generalization to real-world scenarios. To overcome these limitations, we introduce SyncMV4D, the first model that jointly generates synchronized multi-view HOI videos and 4D motions by unifying visual prior, motion dynamics, and multi-view geometry. Our framework features two core innovations: (1) a Multi-view Joint Diffusion (MJD) model that co-generates HOI videos and intermediate motions, and (2) a Diffusion Points Aligner (DPA) that refines the coarse intermediate motion into globally aligned 4D metric point tracks. To tightly couple 2D appearance with 4D dynamics, we establish a closed-loop, mutually enhancing cycle. During the diffusion denoising process, the generated video conditions the refinement of the 4D motion, while the aligned 4D point tracks are reprojected to guide next-step joint generation. Experimentally, our method demonstrates superior performance to state-of-the-art alternatives in visual realism, motion plausibility, and multi-view consistency.

手物交互多视角生成扩散模型4D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。