让语音随人脸表情和情绪强度变化,提升虚拟角色表现力与无障碍体验。
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
- 基于面部图像和情绪强度调节语音情感,无需标注数据训练。
- 在LRS3、CREMA-D、MELD数据集上实现零样本情绪语音合成。
- 适合虚拟角色配音与视障用户听漫画,增强交互沉浸感。
我们提出FEIM-TTS,一种创新的零样本文本到语音(TTS)模型,能够根据面部图像生成情绪化语音,并由情绪强度动态调节。该模型利用深度学习技术解析面部线索,调整情感细微差别,且不依赖标注数据。针对音频-视觉-情感数据稀疏的问题,模型在LRS3、CREMA-D和MELD数据集上进行训练,展现出良好适应性。其独特能力在于生成高质量、与说话人无关的语音,适用于虚拟角色自适应语音生成。此外,通过将情绪细节融入语音,显著提升了视障用户的可访问性,使他们能更完整地享受网络漫画等叙事内容。综合评估证明其在情绪与强度调节方面表现出色,推动了情感语音合成与无障碍技术的发展。示例可访问:https://feim-tts.github.io/。
原文摘要 · Abstract (English)
We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensity. Leveraging deep learning, FEIM-TTS transcends traditional TTS systems by interpreting facial cues and adjusting to emotional nuances without dependence on labeled datasets. To address sparse audio-visual-emotional data, the model is trained using LRS3, CREMA-D, and MELD datasets, demonstrating its adaptability. FEIM-TTS's unique capability to produce high-quality, speaker-agnostic speech makes it suitable for creating adaptable voices for virtual characters. Moreover, FEIM-TTS significantly enhances accessibility for individuals with visual impairments or those who have trouble seeing. By integrating emotional nuances into TTS, our model enables dynamic and engaging auditory experiences for webcomics, allowing visually impaired users to enjoy these narratives more fully. Comprehensive evaluation evidences its proficiency in modulating emotion and intensity, advancing emotional speech synthesis and accessibility. Samples are available at: https://feim-tts.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。