arXiv:2504.01019cs.CV2025-04CVPR被引 13

让多个运动扩散模型动态协作,实现更精细的人体动作生成控制。

MixerMDM: Learnable Composition of Human Motion Diffusion Models

  • 通过对抗训练学习如何动态组合多个预训练运动模型的去噪过程。
  • 可分别控制多人动作细节及整体交互行为,生成质量提升明显。
  • 适合需要精细动作控制的动画、游戏与虚拟人开发场景。

基于文本描述生成人体动作具有挑战性,尤其在需要细粒度控制时。现有方法通过合并多个在不同条件数据集上预训练的运动扩散模型来实现多条件控制,但其组合策略未考虑模型特性与具体文本描述的差异。为此,我们提出MixerMDM,首个可学习的文本条件人体运动扩散模型组合方法。该方法采用对抗式训练,动态学习如何根据输入条件组合各模型的去噪过程。通过融合单人与多人运动扩散模型,MixerMDM实现了对每个个体动作动态及整体交互的细粒度控制。此外,我们提出一种新评估方法,首次在该任务中通过计算混合生成动作与条件间的对齐度,以及MixerMDM在去噪过程中自适应调整混合的能力,来衡量个体质量与交互质量。

原文摘要 · Abstract (English)

Generating human motion guided by conditions such as textual descriptions is challenging due to the need for datasets with pairs of high-quality motion and their corresponding conditions. The difficulty increases when aiming for finer control in the generation. To that end, prior works have proposed to combine several motion diffusion models pre-trained on datasets with different types of conditions, thus allowing control with multiple conditions. However, the proposed merging strategies overlook that the optimal way to combine the generation processes might depend on the particularities of each pre-trained generative model and also the specific textual descriptions. In this context, we introduce MixerMDM, the first learnable model composition technique for combining pre-trained text-conditioned human motion diffusion models. Unlike previous approaches, MixerMDM provides a dynamic mixing strategy that is trained in an adversarial fashion to learn to combine the denoising process of each model depending on the set of conditions driving the generation. By using MixerMDM to combine single- and multi-person motion diffusion models, we achieve fine-grained control on the dynamics of every person individually, and also on the overall interaction. Furthermore, we propose a new evaluation technique that, for the first time in this task, measures the interaction and individual quality by computing the alignment between the mixed generated motions and their conditions as well as the capabilities of MixerMDM to adapt the mixing throughout the denoising process depending on the motions to mix.

动作生成扩散模型动态组合多角色控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。