用音文同步信息生成更自然的全身手势,避免机械感。
ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
- 基于扩散模型,融合音频与文本信息生成手势
- 设计情绪噪声分类器,避免旋律失真并精准控制情感
- 支持音控与文本塑形双模式,适合动画与虚拟人应用
现有手势生成方法主要依赖音频特征生成上半身动作,忽视语义内容、情感和行走运动,导致动作僵硬且无法传达真实语义。本文提出ExpGest,一种利用音文同步信息生成富有表现力的全身手势的新框架。不同于AdaIN或独热编码方法,我们设计了噪声情绪分类器以优化对抗方向噪声,避免旋律失真,并引导结果向指定情绪靠拢。此外,通过在潜在空间中对齐语义与手势,提升了泛化能力。ExpGest是首个支持混合生成模式(音频驱动与文本塑形)的扩散模型手势生成框架。实验表明,该框架能有效学习结合文本驱动运动与音频诱发手势的数据集,初步结果证明其在说话者全局动作的表现力、自然度和可控性上优于当前最优模型。
原文摘要 · Abstract (English)
Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in stiff, mechanical gestures that fail to convey the true meaning of audio content. We introduce ExpGest, a novel framework leveraging synchronized text and audio information to generate expressive full-body gestures. Unlike AdaIN or one-hot encoding methods, we design a noise emotion classifier for optimizing adversarial direction noise, avoiding melody distortion and guiding results towards specified emotions. Moreover, aligning semantic and gestures in the latent space provides better generalization capabilities. ExpGest, a diffusion model-based gesture generation framework, is the first attempt to offer mixed generation modes, including audio-driven gestures and text-shaped motion. Experiments show that our framework effectively learns from combined text-driven motion and audio-induced gesture datasets, and preliminary results demonstrate that ExpGest achieves more expressive, natural, and controllable global motion in speakers compared to state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。