用人脸生成语音,还能通过文字控制语调、节奏等细节。
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
- 结合音视频与纯音频数据训练,提升语音质量。
- 支持真人脸和艺术肖像生成语音,适应性强。
- 通过采样+提示词实现稳定一致的语音合成,适合创意应用。
本文研究多模态可控文本转语音技术,即根据人脸图像生成语音,并通过自然语言描述控制输出语音的语速、噪声水平、距离感、语调、场景等特征。针对当前人脸驱动语音合成系统存在的三个挑战:1)为缓解音视频语料音频质量不足的问题,提出利用高质量纯音频语料进行补充训练;2)为使语音生成不仅限于真实人脸,还可适用于艺术肖像,提出对输入人脸图像进行风格化增强;3)为应对人脸到语音的一对多映射问题并保证生成一致性,提出先采用基于采样的解码方式,再通过生成的语音样本进行提示(prompting)。实验验证了所提模型在人脸驱动语音合成任务中的有效性。
原文摘要 · Abstract (English)
This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with natural text description. Specifically, we aim to mitigate the following three challenges in face-driven TTS systems. 1) To overcome the limited audio quality of audio-visual speech corpora, we propose a training method that additionally utilizes high-quality audio-only speech corpora. 2) To generate voices not only from real human faces but also from artistic portraits, we propose augmenting the input face image with stylization. 3) To consider one-to-many possibilities in face-to-voice mapping and ensure consistent voice generation at the same time, we propose to first employ sampling-based decoding and then use prompting with generated speech samples. Experimental results validate the proposed model's effectiveness in face-driven voice synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。