用分步流匹配提升语音合成的连续声谱建模质量
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
- 基于自回归语言模型与逐标记流匹配结合
- 分层粗到细生成,显著提升语音时序一致性
- 适合追求高质量语音合成的开发者和研究者
为推进连续值标记建模与时间一致性增强,我们提出FELLE,一种将语言建模与逐标记流匹配相结合的自回归模型。通过利用语言模型的自回归特性与流匹配的生成优势,FELLE有效预测连续值标记(梅尔频谱图)。针对每个连续值标记,FELLE通过引入前一步信息来调整流匹配中的通用先验分布,从而提升生成一致性和稳定性。此外,为进一步提升合成质量,FELLE引入粗到细的流匹配机制,分层生成连续值标记,条件依赖于语言模型输出。实验表明,将流匹配技术融入自回归梅尔频谱建模具有显著潜力,大幅提升了语音合成质量,相关结果见 https://aka.ms/felle。
原文摘要 · Abstract (English)
To advance continuous-valued token modeling and temporal-coherence enforcement, we propose FELLE, an autoregressive model that integrates language modeling with token-wise flow matching. By leveraging the autoregressive nature of language models and the generative efficacy of flow matching, FELLE effectively predicts continuous-valued tokens (mel-spectrograms). For each continuous-valued token, FELLE modifies the general prior distribution in flow matching by incorporating information from the previous step, improving coherence and stability. Furthermore, to enhance synthesis quality, FELLE introduces a coarse-to-fine flow-matching mechanism, generating continuous-valued tokens hierarchically, conditioned on the language model's output. Experimental results demonstrate the potential of incorporating flow-matching techniques in autoregressive mel-spectrogram modeling, leading to significant improvements in TTS generation quality, as shown in https://aka.ms/felle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。