提出新方法分离说话人信息,让语音分音节更准确。
Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
- 在固定长度段内回归去噪教师目标,分离说话人干扰
- 音节边界检测与分段聚类达到当前最佳性能
- 适合需要纯净音节表示的语音建模任务
无监督音节分词旨在从原始语音中学习捕捉潜在语言内容结构的离散音节标记。现有方法采用预训练HuBERT的师生蒸馏机制,将隐含语音帧表示组织成音节片段。然而,当使用话语级交叉熵目标训练时,模型会预测说话人身份而非语言内容,损害音节标记的纯净性。为此,我们提出一种说话人解耦的音节分词器,通过在固定长度块内将受说话人影响的学生表示回归到干净的教师目标,以实现更纯净的音节表示。实验表明,该方法在音节边界检测和音节片段聚类上均达到最先进性能。此外,基于该音节标记训练的语音语言模型,在句法和语义理解上比基于音素的SpiRit-LM提升7%相对性能。
原文摘要 · Abstract (English)
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achieves a 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。