arXiv:2409.10103cs.CLcs.SD2024-09中稿 · IEEE SLT 2024被引 5

让语音模型自动发现音节,还能去掉说话人特征干扰。

Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT

  • 用语音增强和帧级训练分离音节与说话人信息
  • 在LibriSpeech上音节分割和音节质量均超越现有方法
  • 适合做无文本语音建模或音系分析的研究者

自监督语音表征学习已成为从未标注音频中提取有意义特征的关键技术。近期进展表明,从与语言单元相关的特征中推导离散符号具有潜力,可在无需文本的情况下支持多种任务训练。特别是,预训练HuBERT的句级自蒸馏(SD-HuBERT)在中间Transformer层的隐含语音帧表示中诱导出音节结构。在SD-HuBERT中,句级表示通过自注意力层使用特殊CLS token从语音帧特征累积而来。然而我们观察到,该CLS token聚合的信息更关联于说话人身份而非语言内容。为此,我们提出一种仅依赖语音的自监督微调方法,将音节单位与说话人信息分离。该方法引入说话人扰动作为数据增强,并采用帧级训练目标,防止CLS token聚合副语言信息。实验结果表明,该方法在Librispeech数据集上的大多数音节分割与音节单元质量指标上优于当前最优方法,证实其在促进纯语音模型音节组织方面的有效性。

原文摘要 · Abstract (English)

Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features correlated with linguistic units, which enables text-less training across diverse tasks. In particular, sentence-level Self-Distillation of the pretrained HuBERT (SD-HuBERT) induces syllabic structures within latent speech frame representations extracted from an intermediate Transformer layer. In SD-HuBERT, sentence-level representation is accumulated from speech frame features through self-attention layers using a special CLS token. However, we observe that the information aggregated in the CLS token correlates more with speaker identity than with linguistic content. To address this, we propose a speech-only self-supervised fine-tuning approach that separates syllabic units from speaker information. Our method introduces speaker perturbation as data augmentation and adopts a frame-level training objective to prevent the CLS token from aggregating paralinguistic information. Experimental results show that our approach surpasses the current state-of-the-art method in most syllable segmentation and syllabic unit quality metrics on Librispeech, underscoring its effectiveness in promoting syllabic organization within speech-only models.

语音处理自监督学习音节发现语音表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。