arXiv:2509.22167eess.AS2025-09被引 17

用语义对齐优化语音合成隐空间,提升音质与可懂度

Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis

  • 在隐空间引入语义对齐正则化,平衡重建与生成质量
  • 在LibriSpeech-PC上实现2.10%错误率与0.64说话人相似度
  • 适合追求高保真语音合成的开发者与研究者

梅尔频谱图在零样本文本到语音合成中被广泛使用,但其固有的冗余导致文本-语音对齐效率低下。基于变分自编码器(VAE)的紧凑隐表示已成为更优替代方案,却面临优化困境:高维隐表示虽提升重建质量与说话人相似性,但损害可懂度;低维隐表示虽改善可懂度,却牺牲重建保真度。为此,我们提出语义对齐隐表示(Semantic-VAE),通过在隐空间引入语义对齐正则化,使高维隐表示有效捕捉语义结构,缓解重建与生成之间的权衡。集成至F5-TTS后,该方法在LibriSpeech-PC上取得2.10%的词错误率(WER)和0.64的说话人相似度,优于基于梅尔频谱图的系统及原生声学VAE基线,同时提升训练效率。演示与代码见:https://zhikangniu.github.io/semantic-vae/

原文摘要 · Abstract (English)

Mel-spectrograms have been widely used in zero-shot text-to-speech (TTS); their inherent redundancy leads to inefficiency in text-speech alignment. Compact VAE-based latent representations have emerged as a stronger alternative but exhibit an optimization dilemma: higher-dimensional latents improve reconstruction quality and speaker similarity but degrade intelligibility, while lower-dimensional latents improve intelligibility at the cost of reconstruction fidelity. To overcome this dilemma, we propose Semantic-VAE, which uses semantic alignment regularization in the latent space. This design alleviates the reconstruction-generation trade-off by capturing semantic structure in high-dimensional latent representations. When integrated into F5-TTS, our method achieves 2.10% WER and 0.64 speaker similarity on LibriSpeech-PC, outperforming mel-based systems and vanilla acoustic VAE baselines with improved training efficiency. Demo and codes: https://zhikangniu.github.io/semantic-vae/

语音合成隐空间建模语义对齐VAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。