arXiv:2507.04554cs.SDcs.CL2025-07中稿 · interspeech 2025被引 1

用荷兰档案电视数据训练出当前最优的荷兰语语音模型。

Self-supervised learning of speech representations with Dutch archival data

  • 用Whisper工具处理噪声广播数据,提升预训练质量。
  • 单语预训练比多语更抗域外数据干扰,效果更稳。
  • 在5.5万小时数据上继续预训练,达到顶尖性能。

本文探讨利用荷兰档案电视广播数据进行自监督语音表征学习,重点使用wav2vec 2.0模型。首先研究预训练对数据质量的要求,发现音乐、噪声和说话人重叠会影响自监督学习收敛与下游微调表现。其次,提出有效预处理策略,通过Whisper和WhisperX将嘈杂广播数据转化为高质量预训练数据。第三,对比单语与多语预训练(数据量相当),表明单语预训练对域外数据更具鲁棒性。最后,通过在55,000小时档案数据上持续预训练wav2vec 2.0 XLS-R模型检查点,实现了荷兰语领域的最新最佳性能大型模型。

原文摘要 · Abstract (English)

This paper explores the use of Dutch archival television broadcast data for self-supervised learning of speech foundation models, specifically wav2vec 2.0. We first study data quality assumptions for pre-training, and show how music, noise and speaker overlap affect SSL convergence and downstream fine-tuning performance. Secondly, we explore effectively pre-processing strategies to convert the noisy broadcast dataset into a qualitative dataset for pre-training, by using Whisper and WhisperX. Thirdly, we compare mono-lingual and multi-lingual pre-training with equivalent amounts of data, and show that mono-lingual pre-training is more robust to out-of-domain data. Lastly, we achieve a state-of-the-art LARGE wav2vec 2.0 model for the Dutch language, by a continuation of pre-training a wav2vec 2.0 XLS-R model checkpoint with our 55k hour archival dataset.

语音模型自监督学习多语种数据清洗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。