用手势控制语音语调,让合成语音更自然有表现力
Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
- 用多模态专家混合架构融合语言与手势信息
- 在PATS数据集上语音自然度和手势同步性均超越基线
- 首个将手部动作用于语音韵律调控的神经合成方法
人类交流天然融合语音与肢体动作,手部手势常与语音韵律协同表达意图、情感和强调。尽管近期文本到语音(TTS)系统已开始引入面部表情或口部动作等多模态线索,但手部手势对韵律的影响仍研究不足。本文提出一种新型多模态TTS框架Gesture2Speech,利用视觉手势线索调节合成语音的韵律。受自信且富有表现力的演讲者手势与语音韵律协调的启发,我们设计了一个多模态专家混合(MoE)架构,动态融合语言内容与手势特征,集成于专用风格提取模块。融合表示作为条件输入基于大语言模型的语音解码器,实现与手部动作时序对齐的韵律调控。此外,我们构建了手势-语音对齐损失,显式建模两者的时间对应关系,确保手势与韵律轮廓的精细同步。在PATS数据集上的评估表明,Gesture2Speech在语音自然度和手势-语音同步性方面均优于现有最优基线。据我们所知,这是首个利用手部手势线索进行神经语音合成中韵律控制的工作。演示样本见 https://research.sri-media-analysis.com/aaai26-beeu-gesture2speech/
原文摘要 · Abstract (English)
Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating multimodal cues such as facial expressions or lip movements, the role of hand gestures in shaping prosody remains largely underexplored. We propose a novel multimodal TTS framework, Gesture2Speech, that leverages visual gesture cues to modulate prosody in synthesized speech. Motivated by the observation that confident and expressive speakers coordinate gestures with vocal prosody, we introduce a multimodal Mixture-of-Experts (MoE) architecture that dynamically fuses linguistic content and gesture features within a dedicated style extraction module. The fused representation conditions an LLM-based speech decoder, enabling prosodic modulation that is temporally aligned with hand movements. We further design a gesture-speech alignment loss that explicitly models their temporal correspondence to ensure fine-grained synchrony between gestures and prosodic contours. Evaluations on the PATS dataset show that Gesture2Speech outperforms state-of-the-art baselines in both speech naturalness and gesture-speech synchrony. To the best of our knowledge, this is the first work to utilize hand gesture cues for prosody control in neural speech synthesis. Demo samples are available at https://research.sri-media-analysis.com/aaai26-beeu-gesture2speech/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。