arXiv:2506.00338cs.CLcs.SD2025-06中稿 · INTERSPEECH 2025被引 37

用清洗后的海量数据提升开源语音模型性能

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

  • 构建可扩展的数据清洗流程,处理网络爬取数据的错误标签和音文错位
  • 获得75种语言共16.6万小时的高质量语音数据,训练出更优模型
  • 新模型在多语言任务中媲美甚至超越工业级顶尖模型,适合研究与应用

Open Whisper-style Speech Models(OWSM)项目已开发出一系列使用学术规模资源的完全开源语音基础模型,但其训练数据仍显不足。本文通过整合大规模网络爬取的YODAS数据集(采用Creative Commons许可),显著扩充数据量。然而,由于YODAS数据质量参差,存在语言标签错误和音文对齐问题,带来挑战。为此,我们设计了一套基于公开工具的可扩展数据清洗流程,最终产出包含75种语言、总计166,000小时语音的高质量数据集。在此基础上,我们训练了新一代OWSM v4模型,结合原有OWSM数据,其在多语言基准测试中表现显著优于前代版本,部分场景下甚至达到或超越如Whisper、MMS等前沿工业模型水平。相关清洗后的YODAS数据、预训练模型及全部脚本将通过ESPnet工具包公开发布。

原文摘要 · Abstract (English)

The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient. This work enhances OWSM by integrating YODAS, a large-scale web-crawled dataset with a Creative Commons license. However, incorporating YODAS is nontrivial due to its wild nature, which introduces challenges such as incorrect language labels and audio-text misalignments. To address this, we develop a scalable data-cleaning pipeline using public toolkits, yielding a dataset with 166,000 hours of speech across 75 languages. Our new series of OWSM v4 models, trained on this curated dataset alongside existing OWSM data, significantly outperform previous versions on multilingual benchmarks. Our models even match or surpass frontier industrial models like Whisper and MMS in multiple scenarios. We will publicly release the cleaned YODAS data, pre-trained models, and all associated scripts via the ESPnet toolkit.

语音模型数据清洗多语言开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。