多模态语音合成框架提升风格控制质量与多样性。
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
- 分三阶段设计,融合人脸与文本信息进行语音控制。
- 在人脸与文本提示任务中均优于单模态基线方法。
- 适合需要高保真与细粒度语音控制的研究者使用。
可控语音合成旨在通过参考输入控制生成语音的风格,参考可为多种模态。现有基于人脸的方法因数据质量限制,鲁棒性与泛化能力不足;而文本提示方法则多样性有限且控制粒度粗。尽管多模态方法试图整合多种模态,但其对完全匹配训练数据的依赖严重制约性能与适用性。本文提出一个三阶段多模态可控语音合成框架:人脸编码器采用监督学习与知识蒸馏以解决泛化问题;文本编码器同时在文本-人脸与文本-语音数据上训练,增强语音多样性。实验表明,该方法在基于人脸和文本提示的语音合成任务中均优于单模态基线,验证了其生成高质量语音的有效性。
原文摘要 · Abstract (English)
Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with robustness and generalization due to data quality constraints, while text prompt methods offer limited diversity and fine-grained control. Although multimodal approaches aim to integrate various modalities, their reliance on fully matched training data significantly constrains their performance and applicability. This paper proposes a 3-stage multimodal controllable speech synthesis framework to address these challenges. For face encoder, we use supervised learning and knowledge distillation to tackle generalization issues. Furthermore, the text encoder is trained on both text-face and text-speech data to enhance the diversity of the generated speech. Experimental results demonstrate that this method outperforms single-modal baseline methods in both face based and text prompt based speech synthesis, highlighting its effectiveness in generating high-quality speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。