用多模态信号生成自然有情感的对话手势,让数字人更像真人。
A conversational gesture synthesis system based on emotions and semantics
- 基于扩散模型,融合文本、语音、情绪和初始动作生成手势。
- 在ZeroEGGS数据集上提升手势真实感与语境契合度。
- 支持情绪插值和合成语音泛化,适合虚拟人研发者使用。
随着大语言模型、语音合成和硬件的进步,当前数字人创作的瓶颈在于生成与文本或语音输入自然匹配的动作。本文提出DeepGesture,一种基于扩散模型的手势生成框架,可依据多模态信号(文本、语音、情绪、初始动作)生成富有表现力的伴随言语手势。在DiffuseStyleGesture基础上引入新架构:采用快速文本转录作为语义条件,并实现情绪引导的无分类器扩散,以支持不同情感状态下的可控生成。通过Unity构建完整渲染管线,基于模型输出的BVH文件进行可视化。在ZeroEGGS数据集上的评估显示,DeepGesture生成的手势在人类相似性和语境适配性上均有提升。系统支持情绪状态间的插值,并展现出对分布外语音(包括合成语音)的泛化能力,为实现全模态、情感感知的数字人迈出关键一步。
原文摘要 · Abstract (English)
Along with the explosion of large language models, improvements in speech synthesis, advancements in hardware, and the evolution of computer graphics, the current bottleneck in creating digital humans lies in generating character movements that correspond naturally to text or speech inputs. In this work, we present DeepGesture, a diffusion-based gesture synthesis framework for generating expressive co-speech gestures conditioned on multimodal signals - text, speech, emotion, and seed motion. Built upon the DiffuseStyleGesture model, DeepGesture introduces novel architectural enhancements that improve semantic alignment and emotional expressiveness in generated gestures. Specifically, we integrate fast text transcriptions as semantic conditioning and implement emotion-guided classifier-free diffusion to support controllable gesture generation across affective states. To visualize results, we implement a full rendering pipeline in Unity based on BVH output from the model. Evaluation on the ZeroEGGS dataset shows that DeepGesture produces gestures with improved human-likeness and contextual appropriateness. Our system supports interpolation between emotional states and demonstrates generalization to out-of-distribution speech, including synthetic voices - marking a step forward toward fully multimodal, emotionally aware digital humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。