arXiv:2507.19359cs.CV2025-07ICCV被引 21

让虚拟人说话时的手势更符合语义,提升自然度

SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning

  • 用语义一致性模块融合语音与手势的深层语义
  • 在两个数据集上优于当前最优方法,用户评测更认可真实感
  • 适合需要自然交互的虚拟主播、教育机器人场景

为虚拟角色生成与语音语义一致的手势是一项挑战。现有研究多聚焦节奏性手势,忽略手势的语义上下文。本文提出一种新方法,通过向量量化变分自编码器学习运动先验,并在第二阶段引入语义相干性与相关性模块,联合语音、文本语义和说话人身份生成语义连贯的手势。实验表明,该方法显著提升手势的真实感与语义一致性。在两个共言手势生成基准上,无论是客观指标还是主观评价,均优于现有最先进方法。模型代码、数据集及预训练模型详见 https://semgesture.github.io/。

原文摘要 · Abstract (English)

Creating a virtual avatar with semantically coherent gestures that are aligned with speech is a challenging task. Existing gesture generation research mainly focused on generating rhythmic beat gestures, neglecting the semantic context of the gestures. In this paper, we propose a novel approach for semantic grounding in co-speech gesture generation that integrates semantic information at both fine-grained and global levels. Our approach starts with learning the motion prior through a vector-quantized variational autoencoder. Built on this model, a second-stage module is applied to automatically generate gestures from speech, text-based semantics and speaker identity that ensures consistency between the semantic relevance of generated gestures and co-occurring speech semantics through semantic coherence and relevance modules. Experimental results demonstrate that our approach enhances the realism and coherence of semantic gestures. Extensive experiments and user studies show that our method outperforms state-of-the-art approaches across two benchmarks in co-speech gesture generation in both objective and subjective metrics. The qualitative results of our model, code, dataset and pre-trained models can be viewed at https://semgesture.github.io/.

手势生成语义对齐虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。