arXiv:2504.09209cs.GRcs.CV2025-04中稿 · ACM Multimedia 202…被引 24

用语音查询注意力机制,精准选择手势关键帧生成自然动作。

EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

  • 通过语音-动作对齐模块构建联合空间,用可学习语音查询定位关键动作帧。
  • 基于语音查询的注意力机制,使模型在高注意力帧上进行选择性掩码。
  • 适合需要高质量语音驱动手势生成的研究者或动画制作人员。

掩码建模在共语手势生成中展现出潜力,但难以准确识别语义重要的帧以实现有效掩码。本文提出一种基于语音查询注意力的掩码建模框架,用于整体共语动作生成。核心思想是利用与动作对齐的语音特征引导掩码过程,有选择地掩码节奏相关和语义表达性强的动作帧。具体而言,首先设计运动-音频对齐模块(MAM),构建潜在的运动-音频联合空间,将低层和高层语音特征投影其中,通过可学习的语音查询实现动作对齐的语音表示。随后引入语音查询注意力机制(SQA),通过运动键与语音查询的交互计算帧级注意力分数,指导对高注意力分数动作帧的选择性掩码。最后,将动作对齐的语音特征注入生成网络,促进共语动作生成。定性和定量评估表明,该方法优于现有最先进方法,能生成高质量共语动作。

原文摘要 · Abstract (English)

Masked modeling has shown promise in co-speech gesture generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we propose a speech-queried attention-based mask modeling framework for holistic co-speech gesture generation. Our key insight is to leverage motion-aligned speech features to guide the masked motion modeling process, selectively masking rhythm-related and semantically expressive motion frames. Specifically, we first propose a motion-audio alignment module (MAM) to construct a latent motion-audio joint space. In this space, both low-level and high-level speech features are projected, enabling motion-aligned speech representation using learnable speech queries. Then, a speech-queried attention mechanism (SQA) is introduced to compute frame-level attention scores through interactions between motion keys and speech queries, guiding selective masking toward motion frames with high attention scores. Finally, the motion-aligned speech features are also injected into the generation network to facilitate co-speech motion generation. Qualitative and quantitative evaluations confirm that our method outperforms existing state-of-the-art approaches, successfully producing high-quality co-speech motion.

手势生成语音驱动注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。