仅用归一化与理论权重即可实现高质量人体动作与形态生成。
Unconditional Human Motion and Shape Generation via Balanced Score-Based Diffusion
- 基于分数扩散模型,仅靠特征归一化和解析权重提升性能。
- 无需额外损失或参数,直接生成动作与身体形态,精度媲美顶尖方法。
- 适合关注生成质量与模型简洁性的研究人员参考。
近期研究探索了多种人体动作生成模型,包括变分自编码器(VAEs)、生成对抗网络(GANs)和基于扩散的模型。尽管方法各异,许多模型依赖过参数化的输入特征和辅助损失来提升实验结果。然而,这些策略对扩散模型而言并非必要。本文表明,仅通过精心设计的特征空间归一化和标准L2分数匹配损失的解析权重,即可在无条件人体动作生成上达到与当前最优水平相当的效果,并直接生成动作与身体形态,避免了事后从关节恢复形状的耗时过程。我们逐步构建方法,每个组件均有明确的理论动机,并通过针对性消融实验证明各部分独立的有效性。
原文摘要 · Abstract (English)
Recent work has explored a range of model families for human motion generation, including Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and diffusion-based models. Despite their differences, many methods rely on over-parameterized input features and auxiliary losses to improve empirical results. These strategies should not be strictly necessary for diffusion models to match the human motion distribution. We show that on par with state-of-the-art results in unconditional human motion generation are achievable with a score-based diffusion model using only careful feature-space normalization and analytically derived weightings for the standard L2 score-matching loss, while generating both motion and shape directly, thereby avoiding slow post hoc shape recovery from joints. We build the method step by step, with a clear theoretical motivation for each component, and provide targeted ablations demonstrating the effectiveness of each proposed addition in isolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。