arXiv:2410.14540cs.CV2024-10被引 4

用多模态扩散模型做人体姿态先验,提升3D姿态估计准确性

Multi-modal Pose Diffuser: A Multimodal Generative Conditional Pose Prior

  • 用多模态条件扩散模型生成符合人体结构的合理姿态
  • 在姿态估计/去噪/补全任务中均优于现有方法
  • 支持图像和文本输入,适合需要上下文理解的应用

SMPL模型在3D人体姿态估计中起关键作用,提供简洁有效的身体表示。然而,在人体网格回归等任务中保持SMPL配置的有效性仍具挑战,亟需一个能区分真实人体姿态的鲁棒姿态先验。为此,我们提出MOPED:首个将多模态条件扩散模型作为SMPL姿态参数先验的方法。该方法具备强大的无条件姿态生成能力,并可基于图像、文本等多模态输入进行条件控制。这种能力通过引入传统姿态先验常忽略的上下文信息,显著提升了适用性。在姿态估计、姿态去噪和姿态补全三项任务上的大量实验表明,基于多模态扩散模型的姿态先验显著优于现有方法,说明其能捕捉更广泛的合理人体姿态。

原文摘要 · Abstract (English)

The Skinned Multi-Person Linear (SMPL) model plays a crucial role in 3D human pose estimation, providing a streamlined yet effective representation of the human body. However, ensuring the validity of SMPL configurations during tasks such as human mesh regression remains a significant challenge , highlighting the necessity for a robust human pose prior capable of discerning realistic human poses. To address this, we introduce MOPED: \underline{M}ulti-m\underline{O}dal \underline{P}os\underline{E} \underline{D}iffuser. MOPED is the first method to leverage a novel multi-modal conditional diffusion model as a prior for SMPL pose parameters. Our method offers powerful unconditional pose generation with the ability to condition on multi-modal inputs such as images and text. This capability enhances the applicability of our approach by incorporating additional context often overlooked in traditional pose priors. Extensive experiments across three distinct tasks-pose estimation, pose denoising, and pose completion-demonstrate that our multi-modal diffusion model-based prior significantly outperforms existing methods. These results indicate that our model captures a broader spectrum of plausible human poses.

姿态生成扩散模型多模态3D人体建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。