分离层次化生成3D人互动,更真实可控。
Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation
- 分层解耦潜空间,分离交互上下文与个体动作
- 对比学习提升物理合理性,避免穿模错位
- 扩散模型结合自适应归一化,生成高质量动作
生成逼真的3D人-人互动(HHI)需要同时建模代理的物理合理性和交互语义。现有方法将所有运动信息压缩到单一潜在表示中,难以捕捉细粒度动作和跨代理交互,常导致语义错位和物理上不合理的现象,如穿模或接触缺失。我们提出解耦层次化变分自编码器(DHVAE),用于结构化且可控的HHI生成。DHVAE通过CoTransformer模块显式地将全局交互上下文与个体运动模式解耦至分离的潜空间。为减少HHI中的不合理接触,引入对比学习约束,增强潜空间的判别性与物理合理性。为实现高保真交互合成,DHVAE在层次潜空间中采用基于DDIM的扩散去噪过程,并使用跳连结构的AdaLN-Transformer去噪器进行优化。大量评估表明,相较于基线方法,DHVAE在运动保真度、文本对齐性和物理合理性方面均表现更优,且计算效率更高。
原文摘要 · Abstract (English)
Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single latent representation, limiting their ability to capture fine-grained actions and inter-agent interactions. This often leads to semantic misalignment and physically implausible artifacts, such as penetration or missed contact. We propose Disentangled Hierarchical Variational Autoencoder (DHVAE) based latent diffusion for structured and controllable HHI generation. DHVAE explicitly disentangles the global interaction context and individual motion patterns into a decoupled latent structure by employing a CoTransformer module. To mitigate implausible and physically inconsistent contacts in HHI, we incorporate contrastive learning constraints with our DHVAE to promote a more discriminative and physically plausible latent interaction space. For high-fidelity interaction synthesis, DHVAE employs a DDIM-based diffusion denoising process in the hierarchical latent space, enhanced by a skip-connected AdaLN-Transformer denoiser. Extensive evaluations show that DHVAE achieves superior motion fidelity, text alignment, and physical plausibility with greater computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。