同步扩散模型实现多人多物交互动作的逼真生成
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
- 用统一扩散模型捕捉多人多物运动联合分布
- 在4个数据集上超越现有最佳方法,动作更自然
- 适合虚拟现实与角色动画领域研究者
在虚拟现实和人物动画中,生成逼真的多人多物交互动作是一个关键问题。不同于常见的一人或一手与单一物体交互的场景,本文研究更具通用性的多体设置,即任意数量的人体、手部与物体同时交互。这种复杂性带来显著挑战,因各身体间存在高度相关性和相互影响,难以保持动作同步。为此,本文提出SyncDiff,一种基于同步运动扩散策略的多体交互生成新方法。SyncDiff采用单一扩散模型捕捉多体运动的联合分布,并引入频域运动分解方案提升动作保真度;同时设计新的对齐评分机制,强调不同身体动作的同步性。通过显式同步策略联合优化数据似然与对齐似然。在四个不同配置的多体场景数据集上进行的大量实验表明,SyncDiff在生成质量上显著优于现有最先进方法。
原文摘要 · Abstract (English)
Synthesizing realistic human-object interaction motions is a critical problem in VR/AR and human animation. Unlike the commonly studied scenarios involving a single human or hand interacting with one object, we address a more generic multi-body setting with arbitrary numbers of humans, hands, and objects. This complexity introduces significant challenges in synchronizing motions due to the high correlations and mutual influences among bodies. To address these challenges, we introduce SyncDiff, a novel method for multi-body interaction synthesis using a synchronized motion diffusion strategy. SyncDiff employs a single diffusion model to capture the joint distribution of multi-body motions. To enhance motion fidelity, we propose a frequency-domain motion decomposition scheme. Additionally, we introduce a new set of alignment scores to emphasize the synchronization of different body motions. SyncDiff jointly optimizes both data sample likelihood and alignment likelihood through an explicit synchronization strategy. Extensive experiments across four datasets with various multi-body configurations demonstrate the superiority of SyncDiff over existing state-of-the-art motion synthesis methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。