arXiv:2508.12001eess.AS2025-08

用专家混合模型提升语音合成的时长建模能力,让发音更自然。

FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis

  • 引入专家混合时长预测器,捕捉音素时长和说话人差异。
  • 在多个数据集上实现更高合成质量与推理效率。
  • 适合追求高自然度语音合成的研究者与开发者。

当前非自回归(NAR)文本到语音(TTS)系统仍难以建模多样且与说话人相关的时长变化。我们进一步发现,更丰富的时长变化会增加基于HiFi-GAN声码器的合成难度,导致频谱伪影和不稳定的时频结构。为此,我们提出FNH-TTS,一种基于VITS的端到端TTS系统,采用专家混合时长建模与鲁棒声码器级合成。具体地,我们引入了混合专家时长预测器(MoE-DP),以捕捉多样化的音素时长模式和说话人特有的语速特征。为将更丰富的时长变化转化为稳定的波形生成,我们还集成了具有协同多带与子带判别器的VOCOS风格声码器。在LJSpeech、VCTK和LibriTTS上的实验表明,FNH-TTS在合成质量、时长类别准确率、声码器重建质量及推理效率方面均有提升。进一步分析显示,MoE-DP是时长建模改进的主要来源,而更强的声码器组件对在丰富时长变化下的鲁棒合成至关重要。

原文摘要 · Abstract (English)

Current non-autoregressive (NAR) text-to-speech (TTS) systems still struggle to model diverse and speaker-dependent duration variation. We further observe that richer duration variation can increase the synthesis difficulty of existing HiFi-GAN-based vocoders, leading to spectral artifacts and unstable time-frequency structures. To address these issues, we propose FNH-TTS, a VITS-based end-to-end TTS system with Mixture-of-Experts duration modeling and robust vocoder-side synthesis. Specifically, we introduce a Mixture-of-Experts Duration Predictor (MoE-DP) to capture diverse phoneme duration patterns and speaker-dependent speaking-rate characteristics. To convert richer duration variation into stable waveform generation, we further integrate a VOCOS-style vocoder with Collaborative Multi-Band and Sub-Band Discriminators. Experiments on LJSpeech, VCTK, and LibriTTS show that FNH-TTS achieves improved synthesis quality, duration-category accuracy, vocoder reconstruction quality, and inference efficiency. Further analysis shows that MoE-DP is the main source of improved duration modeling, while stronger vocoder-side components are necessary for robust synthesis under richer duration variation.

语音合成专家混合时长建模声码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。