arXiv:2511.22863cs.CV2025-11被引 1

用文字描述生成配合说话的手势,让动作更自然有逻辑。

CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation

  • 通过多粒度手势描述生成,填补手势数据缺乏文字标注的空白。
  • 实现语音与手势在节奏和语义上的高度同步,生成质量优于现有方法。
  • 适合做智能交互、虚拟人动效的开发者参考,尤其关注动作自然性。

共说话手势生成显著提升了人机交互能力,但因缺少文本驱动的非即兴手势(如边说话边鞠躬)而受限。现有方法面临两大挑战:一是手势数据集缺乏描述性文字标注导致语义先验缺失;二是难以实现多模态协同控制。本文提出CoordSpeaker框架,首次利用手势理解与描述生成来弥合语义鸿沟,并构建双向手势-文本映射。首先通过运动-语言模型在多粒度上生成描述性手势标题;进而设计统一跨数据集运动表示的条件隐空间扩散模型,结合分层控制去噪器,实现高精度协同手势生成。大量实验表明,该方法生成的手势在节奏上与语音同步,在语义上与任意输入标题一致,且效率与质量均优于现有方法。

原文摘要 · Abstract (English)

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Existing methods face two key challenges: 1) the semantic prior gap due to the lack of descriptive text annotations in gesture datasets, and 2) the difficulty in achieving coordinated multimodal control over gesture generation. To address these challenges, this paper introduces CoordSpeaker, a comprehensive framework that enables coordinated caption-empowered co-speech gesture synthesis. Our approach first bridges the semantic prior gap through a novel gesture captioning framework, leveraging a motion-language model to generate descriptive captions at multiple granularities. Building upon this, we propose a conditional latent diffusion model with unified cross-dataset motion representation and a hierarchically controlled denoiser to achieve highly controlled, coordinated gesture generation. CoordSpeaker pioneers the first exploration of gesture understanding and captioning to tackle the semantic gap in gesture generation while offering a novel perspective of bidirectional gesture-text mapping. Extensive experiments demonstrate that our method produces high-quality gestures that are both rhythmically synchronized with speeches and semantically coherent with arbitrary captions, achieving superior performance with higher efficiency compared to existing approaches.

手势生成文本驱动扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。