融合面部表情的多模态模型提升语音表达自然度
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
- 用面部视觉信息增强预训练语音模型的表达能力
- 情感识别F1值提升5,对话生成更贴近真实情绪
- 适合构建端到端的多模态对话系统开发者
我们提出一种视听语言模型(AVLM),通过将全脸视觉线索融入预训练的表达性语音模型中,实现更具表现力的语音生成。在预训练阶段,探索了多种视觉编码器和多模态融合策略,以确定最优集成方式。后续在情感识别和表达性对话任务上进行微调,相比仅使用语音的基线模型取得显著提升(如情感识别F1提高5)。结果表明,表达性视觉信息对引导语音生成具有重要价值,并为端到端多模态对话系统提供了基础。
原文摘要 · Abstract (English)
We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration approach. Subsequent fine-tuning on emotion recognition and expressive dialogue tasks yields substantial gains over speech-only baselines (e.g., +5 F1 in emotion recognition). AVLM highlights the value of expressive visual information in guiding speech generation and offers a foundation for end-to-end multimodal conversational systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。