让语音合成更懂语义,提升零样本语音生成质量
SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis

- 用语音基础模型引导对齐,优化连续语音表征
- 在Seed-TTS上英文语音识别错误率仅1.71%
- 适合追求高保真零样本语音合成的研究者
连续自回归语音合成最近成为零样本文本到语音(TTS)的有前景方向。然而,现有方法在语义韵律建模与重建驱动的连续语音表示之间存在根本性不匹配,导致模型过度关注低层声学纹理而忽视高层语义连贯性,进一步加剧了自回归生成中的误差累积。为此,我们提出SemaVoice,一种语义感知的连续自回归框架,用于高质量零样本TTS。SemaVoice引入由语音基础模型(SFM)指导的对齐机制,优化连续语音表征,以更好地捕捉局部语义一致性和全局结构关系。这些表征用于条件化自回归框架内的分块扩散头,实现高质量语音合成。在Seed-TTS基准上的实验结果表明,SemaVoice在英语上达到1.71%的单词错误率(WER),并在客观和主观评估中保持与最先进开源系统的竞争力。在固定信息率约束下,不同表示粒度下的显著改进进一步验证了SFM引导对齐的有效性。
原文摘要 · Abstract (English)
Continuous autoregressive speech synthesis has recently emerged as a promising direction for zero-shot text-to-speech (TTS). However, existing methods still suffer from a fundamental mismatch between semantic-prosodic modeling and reconstruction-driven continuous speech representations. This mismatch causes TTS models to focus excessively on low-level acoustic textures at the expense of high-level semantic coherence, further exacerbating error accumulation in autoregressive generation. To address this challenge, we propose SemaVoice, a semantic-aware continuous autoregressive framework for high-fidelity zero-shot TTS. SemaVoice introduces a Speech Foundation Model (SFM) guided alignment mechanism that refines continuous speech representations to better capture both local semantic consistency and global structural relationships. These representations condition a patch-wise diffusion head within the autoregressive framework for high-quality speech synthesis. Experimental results on the Seed-TTS benchmark show that SemaVoice achieves an English WER of 1.71\% and remains highly competitive with state-of-the-art open-source systems in both objective and subjective evaluations. The effectiveness of SFM guided alignment is further confirmed by significant improvements under varying representation granularities with a fixed information-rate constraint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。