arXiv:2510.02848cs.SDcs.AI2025-10

Flamed-TTS实现零样本语音合成,高效生成自然语音且支持动态语速。

Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech

  • 用流匹配重构训练方式,结合离散与连续表示提升语音属性控制。
  • 语音识别错误率仅4%,在自然度、发音相似性等指标上超越现有模型。
  • 无需注意力机制,推理快、计算低,适合实时应用和资源受限场景。

零样本文本到语音(TTS)近年来取得显著进展,使模型能通过简短提示合成语音,模仿说话人身份、语调等特征而无需大量个性化数据。尽管语言模型、扩散模型和流匹配等方法已证明有效,仍存在音素重复、内容误传、推理慢、计算开销大等问题,且时间多样性——影响语音自然度的关键因素——未被充分探索。为此,我们提出Flamed-TTS,一种强调低计算成本、低延迟、高语音保真度及丰富时间多样性的新型零样本TTS框架。通过重构流匹配训练范式,并引入对应语音不同属性的离散与连续表示,实验表明,Flamed-TTS在可懂度、自然度、说话人相似性、声学特征保留和动态语速方面均优于当前最优模型。特别地,其词错误率(WER)仅为4%,显著低于领先基线,同时保持低推理延迟与高语音质量。代码与音频样例可在https://flamed-tts.github.io 获取。

原文摘要 · Abstract (English)

Zero-shot Text-to-Speech (TTS) has recently advanced significantly, enabling models to synthesize speech from text using short, limited-context prompts. These prompts serve as voice exemplars, allowing the model to mimic speaker identity, prosody, and other traits without extensive speaker-specific data. Although recent approaches incorporating language models, diffusion, and flow matching have proven their effectiveness in zero-shot TTS, they still encounter challenges such as unreliable synthesis caused by token repetition or unexpected content transfer, along with slow inference and substantial computational overhead. Moreover, temporal diversity-crucial for enhancing the naturalness of synthesized speech-remains largely underexplored. To address these challenges, we propose Flamed-TTS, a novel zero-shot TTS framework that emphasizes low computational cost, low latency, and high speech fidelity alongside rich temporal diversity. To achieve this, we reformulate the flow matching training paradigm and incorporate both discrete and continuous representations corresponding to different attributes of speech. Experimental results demonstrate that Flamed-TTS surpasses state-of-the-art models in terms of intelligibility, naturalness, speaker similarity, acoustic characteristics preservation, and dynamic pace. Notably, Flamed-TTS achieves the best WER of 4% compared to the leading zero-shot TTS baselines, while maintaining low latency in inference and high fidelity in generated speech. Code and audio samples are available at our demo page https://flamed-tts.github.io.

语音合成零样本流匹配高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。