arXiv:2410.18977cs.CV2024-10被引 21

用注意力机制实现可解释的人体动作生成与精细编辑

Pay Attention and Move Better: Harnessing Attention for Interactive Motion Generation and Training-free Editing

  • 通过自注意力和交叉注意力建模动作帧与文本词的对应关系
  • 仅凭修改注意力图即可实现动作强调、替换等灵活编辑
  • 适合需要可控动作生成与可视化解释的研究者

本研究针对人体动作生成中的交互式编辑问题,提出一种基于注意力机制的动作扩散模型 MotionCLR。传统动作扩散模型缺乏词级文本-动作对应建模能力与可解释性,限制了细粒度编辑。MotionCLR 通过自注意力捕捉帧间序列相似性,调节动作特征顺序;通过交叉注意力建立词与动作时间步的精细对应,激活相关动作片段。基于此,我们设计了无需重新训练的多种编辑方法:动作强调/弱化、原位动作替换、示例驱动生成等。进一步验证了注意力图在动作计数与定位生成方面的可解释潜力。实验表明,该方法在生成质量与编辑灵活性方面表现优异,同时具备良好可解释性。

原文摘要 · Abstract (English)

This research delves into the problem of interactive editing of human motion generation. Previous motion diffusion models lack explicit modeling of the word-level text-motion correspondence and good explainability, hence restricting their fine-grained editing ability. To address this issue, we propose an attention-based motion diffusion model, namely MotionCLR, with CLeaR modeling of attention mechanisms. Technically, MotionCLR models the in-modality and cross-modality interactions with self-attention and cross-attention, respectively. More specifically, the self-attention mechanism aims to measure the sequential similarity between frames and impacts the order of motion features. By contrast, the cross-attention mechanism works to find the fine-grained word-sequence correspondence and activate the corresponding timesteps in the motion sequence. Based on these key properties, we develop a versatile set of simple yet effective motion editing methods via manipulating attention maps, such as motion (de-)emphasizing, in-place motion replacement, and example-based motion generation, etc. For further verification of the explainability of the attention mechanism, we additionally explore the potential of action-counting and grounded motion generation ability via attention maps. Our experimental results show that our method enjoys good generation and editing ability with good explainability.

动作生成注意力机制可解释性编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。