arXiv:2412.16563cs.CV2024-12ICCV被引 43

让说话手势更自然生动,重点动作更突出。

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

  • 分两步生成:先学基础节奏动作,再加关键语义动作
  • 在两个公开数据集上表现优于现有方法,语义表达更丰富
  • 适合需要精准手势动画的虚拟主播、AI助手场景

说话手势生成需兼顾普遍的节奏性动作与稀有的关键语义动作。本文提出SemTalk,通过分阶段学习基础动作与稀疏语义动作,并自适应融合。利用粗到细交叉注意力模块和节奏一致性学习构建与语音节奏同步的基础动作;设计语义强调学习机制,基于帧级语义线索生成语义感知的稀疏动作;最后通过学习的语义得分实现动作融合。在两个公开数据集上的定性和定量对比显示,该方法显著优于当前最优,生成的说话动作兼具稳定基底与增强的语义表现力。

原文摘要 · Abstract (English)

Co-speech gesture generation must carefully integrate common rhythmic motion with rare yet essential semantic gestures. In this work, we propose SemTalk for holistic co-speech gesture generation with frame-level semantic emphasis. Our key insight is to separately learn base motions and sparse motions, and then adaptively fuse them. In particular, coarse2fine cross-attention module and rhythmic consistency learning are explored to establish rhythm-related base motion, ensuring a coherent foundation that synchronizes gestures with the speech rhythm. Subsequently, semantic emphasis learning is designed to generate semantic-aware sparse motion, focusing on frame-level semantic cues. Finally, to integrate sparse motion into the base motion and generate semantic-emphasized co-speech gestures, we further leverage a learned semantic score for adaptive synthesis. Qualitative and quantitative comparisons on two public datasets demonstrate that our method outperforms the state-of-the-art, delivering high-quality co-speech motion with enhanced semantic richness over a stable base motion.

手势生成语义强调语音同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。