arXiv:2604.16135cs.CV2026-04

提出Motion-Adapter,让文本生成复合动作更自然连贯。

Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions

论文配图:Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
图 1 · 摘自论文原文
  • 用解耦注意力图作为结构掩码,指导扩散模型生成复合动作。
  • 在多种提示下生成动作更连贯,超越现有最先进方法。
  • 无需详细描述或大模型,适合生成自然交互动作如边走边挥手。

生成式动作合成近年取得进展,能从多种输入模态生成逼真人体动作。然而,从文本生成包含多个并行动作的复合动作仍面临挑战。现有文本到动作扩散模型存在两大缺陷:(i) 灾难性忽略,因时间信息处理不当导致早期动作被后期动作覆盖;(ii) 注意力坍塌,源于交叉注意力机制中过度特征融合。因此,现有方法常依赖过于详细的文本描述(如“举起右手”)、显式的身体部位指定(如“编辑上半身”)或大型语言模型进行部位解析。这些策略导致物理结构与运动机制语义表征不足,难以实现自然行为(如行走时打招呼)。为此,我们提出Motion-Adapter,一个即插即用模块,通过计算解耦的交叉注意力图,在去噪过程中作为结构掩码引导文本到动作扩散模型。大量实验表明,该方法在多样化文本提示下持续生成更忠实、连贯的复合动作,显著优于当前最先进方法。

原文摘要 · Abstract (English)

Recent advances in generative motion synthesis have enabled the production of realistic human motions from diverse input modalities. However, synthesizing compound actions from texts, which integrate multiple concurrent actions into coherent full-body sequences, remains a major challenge. We identify two key limitations in current text-to-motion diffusion models: (i) catastrophic neglect, where earlier actions are overwritten by later ones due to improper handling of temporal information, and (ii) attention collapse, which arises from excessive feature fusion in cross-attention mechanisms. As a result, existing approaches often depend on overly detailed textual descriptions (e.g., raising right hand), explicit body-part specifications (e.g., editing the upper body), or the use of large language models (LLMs) for body-part interpretation. These strategies lead to deficient semantic representations of physical structures and kinematic mechanisms, limiting the ability to incorporate natural behaviors such as greeting while walking. To address these issues, we propose the Motion-Adapter, a plug-and-play module that guides text-to-motion diffusion models in generating compound actions by computing decoupled cross-attention maps, which serve as structural masks during the denoising process. Extensive experiments demonstrate that our method consistently produces more faithful and coherent compound motions across diverse textual prompts, surpassing state-of-the-art approaches.

动作生成扩散模型文本生成复合动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。