arXiv:2409.10788eess.AScs.SD2024-09被引 2

改进语音预训练的预测目标,提升模型在下游任务的表现。

Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

  • 通过设计更丰富的预测目标,增强模型对语音特征的捕捉能力。
  • 新目标在降噪、内容识别等任务上显著提升性能。
  • 适合研究语音基础模型和自监督学习的学者参考。

语音基础模型(如HuBERT及其变体)在大量无标签语音数据上进行预训练,随后用于多种下游任务。这些模型采用掩码预测目标,即从未掩码上下文中预测被掩码片段的信息。预测目标的选择直接影响模型在下游任务中的表现:捕获语调特征的目标适合说话人相关任务,而捕获音素特征的目标则更适合内容相关任务。此外,预测目标在细节层次上也存在差异——编码精细声学特征的模型在去噪任务中表现更优,而关注高层抽象的目标则在内容任务中更具优势。尽管预测目标至关重要,但其设计选择尚未得到充分研究。本文系统探索了设计选择对下游任务性能的影响,发现当前HuBERT常用的设定可能并非最优。我们提出了构建更丰富预测目标的方法,并通过多项下游任务验证了其有效性,显著提升了模型表现。

原文摘要 · Abstract (English)

Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework impacts their performance on downstream tasks. For instance, models pre-trained with targets that capture prosody learn representations suited for speaker-related tasks, while those pre-trained with targets that capture phonetics learn representations suited for content-related tasks. Moreover, prediction targets can differ in the level of detail they capture. Models pre-trained with targets that encode fine-grained acoustic features perform better on tasks like denoising, while those pre-trained with targets focused on higher-level abstractions are more effective for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks.

语音模型自监督学习预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。