arXiv:2506.16741eess.AScs.AI2025-06中稿 · on Interspeech 202…被引 1

用速度一致性提升语音生成速度,5倍快且音质高

RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching

  • 在流匹配中约束速度场,让少步生成仍保高质量
  • 5步合成即可达高保真,比传统方法快5-10倍
  • 适合对实时性要求高的语音应用

我们提出RapFlow-TTS,一种快速且高保真的文本转语音声学模型,利用流匹配(FM)训练中的速度一致性约束。尽管基于常微分方程(ODE)的语音生成可实现自然语音,但通常需大量生成步骤,造成质量与推理速度的权衡。RapFlow-TTS在FM拉直的ODE轨迹上强制速度场一致性,使少步生成仍保持一致的合成质量。此外,引入时间间隔调度和对抗学习等技术进一步提升少步合成效果。实验表明,该模型在合成步数上较传统FM和基于得分的方法分别减少5倍和10倍,仍能实现高保真语音合成。

原文摘要 · Abstract (English)

We introduce RapFlow-TTS, a rapid and high-fidelity TTS acoustic model that leverages velocity consistency constraints in flow matching (FM) training. Although ordinary differential equation (ODE)-based TTS generation achieves natural-quality speech, it typically requires a large number of generation steps, resulting in a trade-off between quality and inference speed. To address this challenge, RapFlow-TTS enforces consistency in the velocity field along the FM-straightened ODE trajectory, enabling consistent synthetic quality with fewer generation steps. Additionally, we introduce techniques such as time interval scheduling and adversarial learning to further enhance the quality of the few-step synthesis. Experimental results show that RapFlow-TTS achieves high-fidelity speech synthesis with a 5- and 10-fold reduction in synthesis steps than the conventional FM- and score-based approaches, respectively.

语音合成流匹配快速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。