arXiv:2603.26747cs.CVcs.LG2026-03被引 1

用流模型替代扩散模型,让动作生成更快更稳。

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

  • 在相同框架下比较扩散与流模型的生成效果。
  • 流模型训练更快,50轮内达最优,采样步数少3倍。
  • 适合追求高效生成的开发者和动画制作人员。

近期文本驱动的动作生成方法分为离散标记与连续潜在空间两类。MotionGPT3 属于后者,结合学习到的连续动作潜在空间与基于扩散的先验进行文本条件生成。尽管修正流目标在图像和音频生成中已展现出优于扩散模型的收敛性与推理效率,但其在动作生成中的适用性仍不明确。本文在 MotionGPT3 框架内开展受控实证研究,固定模型结构、训练协议与评估设置,仅对比扩散与修正流目标的影响。在 HumanML3D 数据集上的实验表明,修正流在更少训练轮次内收敛,早期即达到强性能,且在相同条件下运动质量不低于或超过扩散模型。此外,流模型在多种推理步数下表现稳定,以更少采样步数实现可比质量,显著提升效率-质量权衡。结果表明,修正流的多项优势可迁移至连续潜在空间的文本到动作生成,强调了生成目标选择对动作先验的重要性。

原文摘要 · Abstract (English)

Recent text-driven motion generation methods span both discrete token-based approaches and continuous-latent formulations. MotionGPT3 exemplifies the latter paradigm, combining a learned continuous motion latent space with a diffusion-based prior for text-conditioned synthesis. While rectified flow objectives have recently demonstrated favorable convergence and inference-time properties relative to diffusion in image and audio generation, it remains unclear whether these advantages transfer cleanly to the motion generation setting. In this work, we conduct a controlled empirical study comparing diffusion and rectified flow objectives within the MotionGPT3 framework. By holding the model architecture, training protocol, and evaluation setup fixed, we isolate the effect of the generative objective on training dynamics, final performance, and inference efficiency. Experiments on the HumanML3D dataset show that rectified flow converges in fewer training epochs, reaches strong test performance earlier, and matches or exceeds diffusion-based motion quality under identical conditions. Moreover, flow-based priors exhibit stable behavior across a wide range of inference step counts and achieve competitive quality with fewer sampling steps, yielding improved efficiency-quality trade-offs. Overall, our results suggest that several known benefits of rectified flow objectives do extend to continuous-latent text-to-motion generation, highlighting the importance of the training objective choice in motion priors.

动作生成流模型高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。