通过多模态输入实现零样本语音合成的风格与说话人控制
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
- 用统一编码器融合文本、音频和音色参考,实现零样本风格控制
- 采用分层Conformer结构融合控制特征,提升语音自然度
- 适合需要个性化语音生成的开发者与研究者使用
我们提出StyleFusion-TTS,一种基于提示和/或音频参考的零样本文本到语音合成系统,可同时控制语音风格与说话人特征。该系统设计了一个通用前端编码器,以紧凑高效的方式利用文本提示、音频参考和说话人音色参考,在完全零样本条件下生成解耦的风格与说话人控制嵌入。此外,提出一种分层Conformer结构,用于融合风格与说话人控制嵌入,旨在在当前先进TTS架构中实现最优特征融合。通过多种主客观指标评估,系统在各项测试中表现优异,表明其在零样本语音合成领域具有重要潜力。
原文摘要 · Abstract (English)
We introduce StyleFusion-TTS, a prompt and/or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research literature. We propose a general front-end encoder as a compact and effective module to utilize multimodal inputs including text prompts, audio references, and speaker timbre references in a fully zero-shot manner and produce disentangled style and speaker control embeddings. Our novel approach also leverages a hierarchical conformer structure for the fusion of style and speaker control embeddings, aiming to achieve optimal feature fusion within the current advanced TTS architecture. StyleFusion-TTS is evaluated through multiple metrics, both subjectively and objectively. The system shows promising performance across our evaluations, suggesting its potential to contribute to the advancement of the field of zero-shot text-to-speech synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。