arXiv:2603.29339cs.SDeess.AS2026-03被引 6

直接在波形潜在空间生成语音,实现高保真零样本克隆。

LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space

  • 跳过中间表示,直接在波形潜在空间用扩散模型生成语音。
  • 在Seed-ZH和Seed-Hard上语音相似度分别达0.818和0.797,超越此前SOTA。
  • 揭示了波形自编码器重建精度越高,语音合成效果不一定越好。

我们提出LongCat-AudioDiT,一种基于扩散模型的非自回归文本转语音(TTS)新方法,达到当前最优性能。不同于以往依赖梅尔频谱等中间声学表征的方法,其核心创新在于直接在波形潜在空间中操作,有效避免误差累积,大幅简化流程,仅需一个波形变分自编码器(Wav-VAE)和扩散主干网络。此外,我们改进推理过程:首先识别并修正长期存在的训练-推理不匹配问题;其次用自适应投影引导替代传统无分类器引导,提升生成质量。实验表明,即使无需复杂多阶段训练或高质量人工标注数据,LongCat-AudioDiT在Seed基准上仍实现当前最优的零样本语音克隆性能,且保持良好可懂性。具体而言,其最大版本LongCat-AudioDiT-3.5B在Seed-ZH上将说话人相似度(SIM)从0.809提升至0.818,在Seed-Hard上从0.776提升至0.797。通过系统消融与分析,验证了所提模块有效性。值得注意的是,我们发现波形自编码器的重建保真度越高,整体语音合成性能反而不一定更好,这一反直觉现象为后续研究提供新视角。代码与模型权重已公开,以促进语音领域进一步研究。

原文摘要 · Abstract (English)

We present LongCat-AudioDiT, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance. Unlike previous methods that rely on intermediate acoustic representations such as mel-spectrograms, the core innovation of LongCat-AudioDiT lies in operating directly within the waveform latent space. This approach effectively mitigates compounding errors and drastically simplifies the TTS pipeline, requiring only a waveform variational autoencoder (Wav-VAE) and a diffusion backbone. Furthermore, we introduce two critical improvements to the inference process: first, we identify and rectify a long-standing training-inference mismatch; second, we replace traditional classifier-free guidance with adaptive projection guidance to elevate generation quality. Experimental results demonstrate that, despite the absence of complex multi-stage training pipelines or high-quality human-annotated datasets, LongCat-AudioDiT achieves SOTA zero-shot voice cloning performance on the Seed benchmark while maintaining competitive intelligibility. Specifically, our largest variant, LongCat-AudioDiT-3.5B, outperforms the previous SOTA model (Seed-TTS), improving the speaker similarity (SIM) scores from 0.809 to 0.818 on Seed-ZH, and from 0.776 to 0.797 on Seed-Hard. Finally, through comprehensive ablation studies and systematic analysis, we validate the effectiveness of our proposed modules. Notably, we investigate the interplay between the Wav-VAE and the TTS backbone, revealing the counterintuitive finding that superior reconstruction fidelity in the Wav-VAE does not necessarily lead to better overall TTS performance. Code and model weights are released to foster further research within the speech community.

语音合成扩散模型零样本克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。