用广播音频训练语音模型,提升多任务表现并揭示数据去重重要性
Data Selection Effects on Self-Supervised Learning of Audio Representations for French Audiovisual Broadcasts
- 构建多样化广播音频预训练数据集,覆盖语音与音乐
- 在多任务上表现优于仅用语音训练的模型
- 通过成员推断攻击揭示数据冗余风险,适合语音与音乐研究者
自监督学习(SSL)语音编码器模型已广泛应用于各类任务。现有模型多在清晰分割的语音数据集(如LibriSpeech)上训练。本文构建了一个大规模、多样化的电视与广播音频预训练语料库,并通过自动工具进行标注。基于这些标注,生成多个子集用于训练音频SSL模型。随后在自动语音识别、语音活动检测、音乐检测及说话人识别等下游任务上评估模型性能。结果表明,在非限定于语音的多样化音频内容上预训练,可显著提升模型表现。此外,通过成员推断攻击评估编码器对训练数据的记忆能力,凸显数据去重的重要性。该统一训练范式有助于弥合语音与音乐机器学习领域的鸿沟。
原文摘要 · Abstract (English)
Audio and speech self-supervised encoder models are now widely used for a lot of different tasks. Many of these models are often trained on clean segmented speech content such as LibriSpeech. In this paper, we look into how the pretraining datasets of such SSL (Self-Supervised Learning) models impact their downstream results. We build a large pretraining corpus of highly diverse TV and Radio broadcast audio content, which we describe with automatic tools. We use these annotations to build smaller subsets, which we use to train audio SSL models. Then, we evaluate the models on multiple downstream tasks such as automatic speech recognition, voice activity and music detection, or speaker recognition. The results show the potential of pretraining SSL models on diverse audio content without restricting it to speech. We also perform a membership inference attack to evaluate the encoder ability to memorize their training datasets, which highlight the importance of data deduplication. This unified training could bridge speech and music machine learning communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。