arXiv:2607.06027cs.SD2026-07

用语音表示的弗雷歇距离提升少步语音合成的可懂度

Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis

  • 训练时引入语音表示的弗雷歇距离正则化,匹配高质量语音统计特征
  • 在四步生成下使英文语音识别错误率从2.23%降至1.41%,相对降低36.5%
  • 无需判别器或推理时计算,适合部署在低延迟场景的语音合成系统

少步扩散与流匹配语音合成模型通常使用局部目标函数进行训练,如条件流匹配、重构和停顿预测。这些损失虽能提供稳定优化,但无法保证生成语音是否符合高质量语音的分布。本文提出语音表示弗雷歇距离损失(SR-FD),作为无分词器流匹配自回归语音合成的训练时分布正则化方法。微调阶段,模型使用与部署时相同的少步采样器生成语音,SR-FD将生成语音的冻结Whisper和CTC特征均值与协方差,匹配来自三个互补内容目标的离线参考统计量。该损失无需判别器,也无需推理时计算。在Seed-TTS英文数据集上,四步SR-FD微调将原始四步VoxCPM2基线的词错误率从2.2279%降至1.4147%,相对降低36.5%,并优于原始十步基线的1.7366%;两项提升在句级配对自助法检验下均显著。说话人相似性与客观质量指标保持在十步水平,错误分析表明增益源于所有提示长度下的内容替换。因此,SR-FD是少步语音合成中提升可懂度的分布正则化方法。

原文摘要 · Abstract (English)

Few-step diffusion and flow-matching text-to-speech (TTS) models are usually trained with local objectives, such as conditional flow matching, reconstruction, and stop prediction. These losses provide stable optimization, but they never ask whether sampled speech follows the distribution of high-quality speech. We propose Speech Representation Fr'echet Distance loss (SR-FD), a training-time distributional regularizer for tokenizer-free flow-matching autoregressive TTS. During fine-tuning, the model synthesizes speech with the same few-step sampler used at deployment, and SR-FD matches the mean and covariance of frozen Whisper and CTC features of this speech to reference statistics computed offline from three complementary content targets. The loss requires no discriminator and no inference-time computation. On Seed-TTS English, four-step SR-FD fine-tuning reduces WER from the original four-step VoxCPM2 baseline's 2.2279% to 1.4147%, a 36.5% relative reduction, and also surpasses the original ten-step baseline at 1.7366%; both gains are significant under an utterance-level paired bootstrap. Speaker similarity and objective quality proxies are preserved at the ten-step level, and an error analysis shows the gain comes from content substitutions across all prompt lengths. SR-FD is thus an intelligibility-improving distributional regularizer for few-step TTS.

语音合成扩散模型可懂度分布匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。