arXiv:2608.07462eess.AScs.SD2026-08

用语义标记指导连续语音生成,提升内容准确性。

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

论文配图:SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
图 1 · 摘自论文原文
  • 训练时用离散语义标记监督连续生成过程
  • 在多任务测试中降低词错率与字符错率
  • 适合追求高内容保真的语音合成研究者

连续潜空间自回归语音生成因其避免量化损失并保留更丰富的声学信息而成为有前景的替代方案。然而,连续声学目标并未显式暴露语言结构,导致自回归语言模型需通过声学预测间接学习语言结构,可能损害生成语音的内容保真度。我们提出 SemBridge,一种仅用于训练的语义标记锚定框架,利用离散语义标记直接监督自回归语言模型状态,并采用语义对齐的声学变分自编码器(Semantic-Aligned Acoustic VAE)在相同语义参考下组织连续目标空间。语义监督仅在训练阶段使用,推理过程仍保持完全连续。我们在零样本文本到语音(TTS)和分数控制的歌唱语音合成(SVS)上评估 SemBridge。在多个基准测试中,该方法在保持竞争力的说话人相似度和感知质量的同时,显著提升了内容准确性,表现为词错误率(WER)和字符错误率(CER)降低。实验表明,对自回归状态学习施加显式语义标记监督是连续语音生成的有效且通用方向。语音样例已提供,模型代码与检查点将公开于 https://github.com/ASLP-lab/SemBridge。

原文摘要 · Abstract (English)

Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit token-level prediction tar- gets. Consequently, the autoregressive language model (LM) must acquire linguistic structure indirectly through acous- tic prediction, which can compromise the content fidelity of generated speech. We propose SemBridge, a training-only semantic-token anchoring framework for continuous-latent autoregressive speech generation. SemBridge uses discrete se- mantic tokens to directly supervise autoregressive LM states and employs a Semantic-Aligned Acoustic VAE to organize the continuous target space under the same semantic refer- ence. The semantic supervision is used only during train- ing, while inference remains entirely continuous. We evalu- ate SemBridge on zero-shot text-to-speech (TTS) and score- conditioned singing voice synthesis (SVS). Across multi- ple benchmarks, SemBridge improves content accuracy, as measured by word and character error rates (WER/CER), while maintaining competitive speaker similarity and percep- tual quality. Experimental results demonstrate that explicit semantic-token supervision for autoregressive state learning is an effective and general direction for continuous speech generation. Speech samples are available.1 The model code and checkpoints will be available at https://github.com/ASLP- lab/SemBridge

语音生成连续潜空间语义监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。