arXiv:2608.16235eess.AS2026-08

通过迭代优化,让语音语义令牌更纯净、更像文本。

Speaker-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement

论文配图:Speaker-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement
图 1 · 摘自论文原文
  • 交替训练语音转单元和文本转单元模型,逐步净化语义令牌。
  • 改进后令牌在跨说话人一致性上提升,语音信息减少。
  • 适合语音转换和语音合成任务,保持高可懂度。

语义语音令牌应在保留语言内容的同时,抑制源自声学输入的说话人和时长相关变化。我们提出迭代语义令牌净化(ISTP),一种由文本可预测性引导的语音到单元(S2U)与文本到单元(T2U)交替训练流程。从初始S2U分词器开始,每轮训练基于去重后的令牌序列的T2U模型;解码得到的T2U预测作为新初始化的S2U模型的连接时序分类目标,其输出又用于监督下一轮T2U模型。该循环逐步对齐两个生成器,并使令牌空间偏向于仅从文本可恢复的信息。在中文和英文上的实验表明,S2U与T2U的一致性显著提升。独立训练的逆分词器进一步显示,优化后的S2U与T2U令牌足以支持高可懂度的语音转换和文本到语音合成。在语音转换中,生成的发音速率更贴近参考源;优化后的令牌也展现出显著更高的跨说话人一致性与更低的探测可恢复说话人信息。

原文摘要 · Abstract (English)

Semantic speech tokens should preserve linguistic content while suppressing speaker- and duration-dependent variation inherited from acoustic inputs. We propose Iterative Semantic Token Purification (ISTP), an alternating speech-to-unit (S2U) and text-to-unit (T2U) training procedure guided by text predictability. Starting from an initial S2U tokenizer, each iteration trains a T2U model on its deduplicated token sequences. The decoded T2U predictions then serve as connectionist temporal classification targets for a newly initialized S2U model, whose outputs supervise the next T2U model. This cycle progressively aligns the two token generators and biases the token space toward information recoverable from text. Experiments on Mandarin and English show substantially improved S2U--T2U agreement. Independently trained de-tokenizers further show that the refined S2U and T2U tokens retain sufficient content for high-intelligibility voice conversion and text-to-speech synthesis. In voice conversion, the generated speaking rate follows the reference more closely. The refined tokens also exhibit substantially improved cross-speaker consistency and reduced probe-recoverable speaker information.

语音语义语音转换令牌净化多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。