arXiv:2512.04499cs.CVcs.GR2025-12

探究运动表示对扩散模型生成质量的影响,发现关键选择决定生成效果。

Back to Basics: Motion Representation Matters for Human Motion Generation Using Diffusion Model

  • 用加权运动数据与噪声之和作为预测目标,改进扩散模型训练
  • 六种运动表示中,不同方案在质量和多样性上差异明显
  • 实验揭示配置选择对训练速度和生成效果的关键影响

扩散模型已成为人体动作合成中广泛使用且成效显著的方法。面向任务的扩散模型显著推动了动作到动作、文本到动作及音频到动作的应用进展。本文通过受控实验,系统研究运动表示与损失函数的基础问题,并梳理生成式运动扩散模型工作流中的各项决策影响。基于代理运动扩散模型(MDM),我们采用以v为预测目标的设定(vMDM),其中v是运动数据与噪声的加权和。旨在深化对潜在数据分布的理解,为提升条件运动扩散模型奠定基础。首先,评估文献中六种常见运动表示在质量与多样性指标上的表现;其次,对比不同配置下的训练时间,探讨加速训练的有效路径;最后,在大规模运动数据集上进行评估分析。实验结果表明,不同运动表示在多种数据集上表现差异显著,且配置选择对模型训练与生成结果具有显著影响,凸显了这些设计决策的重要性与有效性。

原文摘要 · Abstract (English)

Diffusion models have emerged as a widely utilized and successful methodology in human motion synthesis. Task-oriented diffusion models have significantly advanced action-to-motion, text-to-motion, and audio-to-motion applications. In this paper, we investigate fundamental questions regarding motion representations and loss functions in a controlled study, and we enumerate the impacts of various decisions in the workflow of the generative motion diffusion model. To answer these questions, we conduct empirical studies based on a proxy motion diffusion model (MDM). We apply v loss as the prediction objective on MDM (vMDM), where v is the weighted sum of motion data and noise. We aim to enhance the understanding of latent data distributions and provide a foundation for improving the state of conditional motion diffusion models. First, we evaluate the six common motion representations in the literature and compare their performance in terms of quality and diversity metrics. Second, we compare the training time under various configurations to shed light on how to speed up the training process of motion diffusion models. Finally, we also conduct evaluation analysis on a large motion dataset. The results of our experiments indicate clear performance differences across motion representations in diverse datasets. Our results also demonstrate the impacts of distinct configurations on model training and suggest the importance and effectiveness of these decisions on the outcomes of motion diffusion models.

运动生成扩散模型表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。