arXiv:2411.16575cs.CV2024-11CVPR被引 86

改进扩散模型生成动作,解决离散化信息损失问题

Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression

  • 用掩码自回归优化扩散模型,保留连续空间优势
  • 新表示方法使生成动作更自然,多样性提升显著
  • 适合追求高质量动作生成的研究者与开发者

自2023年以来,基于向量量化(VQ)的离散生成方法在人体动作生成中迅速占据主导地位,性能普遍优于扩散模型。然而,将连续动作数据映射为有限离散符号会导致信息丢失,降低生成动作多样性,并限制其作为运动先验或生成引导的能力。相比之下,扩散模型的连续空间特性更有利于克服这些缺陷,且具备更强的可扩展性。本文系统分析了当前VQ方法表现优异的原因,并从动作数据表示与分布角度揭示现有扩散方法的局限。基于此,我们保留扩散模型的核心优势,借鉴VQ方法思想进行渐进式优化,提出一种支持掩码自回归的人体动作扩散模型,采用重构的数据表示与分布。同时,设计更稳健的评估方法。在多个数据集上的实验表明,该方法超越先前方法,达到当前最优性能。

原文摘要 · Abstract (English)

Since 2023, Vector Quantization (VQ)-based discrete generation methods have rapidly dominated human motion generation, primarily surpassing diffusion-based continuous generation methods in standard performance metrics. However, VQ-based methods have inherent limitations. Representing continuous motion data as limited discrete tokens leads to inevitable information loss, reduces the diversity of generated motions, and restricts their ability to function effectively as motion priors or generation guidance. In contrast, the continuous space generation nature of diffusion-based methods makes them well-suited to address these limitations and with even potential for model scalability. In this work, we systematically investigate why current VQ-based methods perform well and explore the limitations of existing diffusion-based methods from the perspective of motion data representation and distribution. Drawing on these insights, we preserve the inherent strengths of a diffusion-based human motion generation model and gradually optimize it with inspiration from VQ-based approaches. Our approach introduces a human motion diffusion model enabled to perform masked autoregression, optimized with a reformed data representation and distribution. Additionally, we propose a more robust evaluation method to assess different approaches. Extensive experiments on various datasets demonstrate our method outperforms previous methods and achieves state-of-the-art performances.

动作生成扩散模型自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。