arXiv:2503.13859cs.CV2025-03

用稀疏关键帧提升动作扩散模型效率与真实感

Less is More: Improving Motion Diffusion Models with Sparse Keyframes

  • 以稀疏关键帧为结构,动态优化生成过程
  • 在更少扩散步骤下实现更好文本对齐和动作真实感
  • 适合需要高效生成的动画、游戏开发场景

近期动作扩散模型在文本到动作生成等任务中取得显著进展,但现有方法将动作表示为密集帧序列,需处理大量冗余或低信息量帧,导致训练复杂度高,尤其在大型动作数据集上受限严重。受专业动画师仅关注稀疏关键帧的启发,我们提出一种基于稀疏几何意义关键帧的新扩散框架。通过掩码非关键帧并高效插值缺失帧,减少计算量;推理时动态优化关键帧掩码,优先保留后期扩散步骤中的关键信息。大量实验表明,该方法在文本对齐与动作真实感方面持续优于当前最优方法,同时在显著更少的扩散步数下仍保持高性能。进一步验证了其作为生成先验的鲁棒性,并可适配多种下游任务。

原文摘要 · Abstract (English)

Recent advances in motion diffusion models have led to remarkable progress in diverse motion generation tasks, including text-to-motion synthesis. However, existing approaches represent motions as dense frame sequences, requiring the model to process redundant or less informative frames. The processing of dense animation frames imposes significant training complexity, especially when learning intricate distributions of large motion datasets even with modern neural architectures. This severely limits the performance of generative motion models for downstream tasks. Inspired by professional animators who mainly focus on sparse keyframes, we propose a novel diffusion framework explicitly designed around sparse and geometrically meaningful keyframes. Our method reduces computation by masking non-keyframes and efficiently interpolating missing frames. We dynamically refine the keyframe mask during inference to prioritize informative frames in later diffusion steps. Extensive experiments show that our approach consistently outperforms state-of-the-art methods in text alignment and motion realism, while also effectively maintaining high performance at significantly fewer diffusion steps. We further validate the robustness of our framework by using it as a generative prior and adapting it to different downstream tasks.

动作生成扩散模型稀疏建模关键帧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。