用视频提升语音合成的语气表现力
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video
- 融合文本与视频信息预测语音语调
- 实验显示加入视觉特征后语音更生动
- 适合做语音合成与多模态研究者
文本到语音(TTS)合成面临同一文本生成不同语调语音的挑战。以往研究通过从文本和语音中预测语调信息来解决,但许多应用场景中可用的视频信息仍被忽视。本文探讨了引入视觉上下文对语调预测的潜力,提出新模型VisualSpeech,结合视觉与文本信息以改进TTS中的语调生成。实验证明,融入视觉特征能有效提升语调建模能力,使合成语音更具表现力。音频样本可访问:https://ariameetgit.github.io/VISUALSPEECH-SAMPLES/
原文摘要 · Abstract (English)
Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text and speech, additional contextual information, such as video, remains under-utilized despite being available in many applications. This paper investigates the potential of integrating visual context to enhance prosody prediction. We propose a novel model, VisualSpeech, which incorporates visual and textual information for improving prosody generation in TTS. Empirical results indicate that incorporating visual features improves prosodic modeling, enhancing the expressiveness of the synthesized speech. Audio samples are available at https://ariameetgit.github.io/VISUALSPEECH-SAMPLES/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。