arXiv:2509.09631cs.SDcs.CL2025-09中稿 · Interspeech 2026

用离散流匹配实现低延迟零样本语音合成

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching

  • 基于离散流匹配构建语音生成框架,避免连续空间优化难题
  • 无需自回归生成,推理速度显著提升,支持零样本语音克隆
  • 适合对实时性要求高的语音应用,如智能助手、语音交互

零样本语音合成在模仿未见嗓音方面取得显著进展,但生成质量与推理效率的平衡仍具挑战。自回归模型存在高延迟问题,而基于扩散的方法受限于训练时配置。此外,多数基于流的方法在连续空间中运行,因连续标记空间比离散空间更复杂,带来优化困难。为此,我们提出 DiFlow-TTS,一种基于离散流匹配的新型零样本语音合成框架。模型包含一个确定性的音素-内容映射器用于语言建模,以及一个因子化离散流去噪器,可同时生成语调和声学标记序列。实验结果表明,该方法在多项评估指标上均表现优异。

原文摘要 · Abstract (English)

Zero-shot text-to-speech (TTS) has made significant progress in replicating unseen voices, yet balancing generation quality and inference efficiency remains challenging. Autoregressive models suffer from high latency, while diffusion-based approaches are constrained by training-time configurations. Moreover, most flow-based methods operate in continuous space, which introduces optimization challenges because continuous token spaces are inherently more complex than discrete ones. To address these limitations, we propose DiFlow-TTS, a novel zero-shot TTS framework based on discrete flow matching. The model consists of a deterministic Phoneme-Content Mapper for linguistic modeling and a Factorized Discrete Flow Denoiser that simultaneously generates prosody and acoustic token streams. Experimental results demonstrate the effectiveness of our approach across multiple evaluation metrics.

语音合成离散流零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。