统一生成多视角手物交互视频与3D轨迹,实现外观与运动同步
HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

- 用多视图扩散变换器联合建模2D图像与3D点轨迹
- 在多个数据集上达到最高视觉质量与几何一致性
- 适合动画生成与具身智能中的多视角交互合成
手物交互(HOI)合成是动画制作和具身AI的核心。尽管视频基础模型具备强先验,但复杂手部动作与遮挡仍使多视角一致的HOI合成极具挑战。本文提出HarmoHOI,一种统一的扩散框架,可协同生成同步的多视角HOI视频与全局对齐的3D点轨迹。核心洞察是:鲁棒的多视角一致性需依赖全局对齐的3D几何与运动。为此,我们设计了多视图扩散变换器,将点轨迹表示为伪视频,实现3D几何信号与基础模型2D潜在空间的对齐,有效缩小域差距并利用先验。进一步提出全局运动对齐扩散,将粗略点轨迹优化为度量尺度、全局对齐的3D轨迹。HarmoHOI支持去噪过程中2D外观与3D运动的动态协同演化。为缓解多视角HOI数据稀缺,采用混合数据课程学习策略,成功将单视角通用先验迁移至同步多视角生成。实验表明,HarmoHOI在视觉质量、动作合理性及多视角几何一致性上均达到当前最优。项目页面见https://droliven.github.io/HarmoHOI_project。
原文摘要 · Abstract (English)
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions and occlusions. We present HarmoHOI, a unified diffusion framework that jointly and harmoniously generates synchronized multi-view HOI videos and globally aligned 3D point tracks. Our core insight is that robust multi-view consistency fundamentally requires globally aligned 3D geometry and motion. To this end, we propose a Mixture of Multi-view Diffusion Transformer that co-models RGB videos and 3D point tracks. By representing point tracks as pseudo-videos, we align 3D geometric signals with the 2D latent space of foundation models, thereby minimizing the domain gap and easing adaptation of priors. To further ensure geometry consistency, we introduce Global Motion Aligning Diffusion, which refines coarse point tracks into metric-scale, globally aligned 3D trajectories. HarmoHOI enables on-the-fly co-evolution of 2D appearance and 3D motion during denoising. To overcome the scarcity of multi-view HOI data, we employ a hybrid data curriculum learning strategy that successfully transfers generic priors from single-view data to synchronized multi-view generation. Experimental results show that HarmoHOI achieves state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency. Project page available at https://droliven.github.io/HarmoHOI_project.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。