arXiv:2410.01036cs.CLcs.AI2024-10EMNLP被引 16

开源语音模型训练数据突破95万小时,覆盖欧盟24种语言

MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages

论文配图:MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
图 1 · 摘自论文原文
  • 收集95万小时欧盟语言语音数据,符合开源许可要求
  • 提供44.1万小时未标注数据的自动转录文本
  • 助力构建完全开源的欧盟语言语音基础模型

随着基础模型(FMs)兴起及监管对风险与影响的关注,开源模型备受重视。然而现有语音基础模型(SFMs)虽宣称开源,实际仍缺乏公开模型权重、代码和训练数据,无法真正满足开源原则。本文首次填补这一空白,聚焦欧盟24种官方语言,通过调研符合开源许可的语音识别数据集和无标签语音语料库,共收集95万小时训练数据。此外,我们以宽松的CC-BY许可发布其中44.1万小时无标签数据的自动转录结果,为欧盟语言的完全开源语音基础模型构建提供支持。

原文摘要 · Abstract (English)

The rise of foundation models (FMs), coupled with regulatory efforts addressing their risks and impacts, has sparked significant interest in open-source models. However, existing speech FMs (SFMs) fall short of full compliance with the open-source principles, even if claimed otherwise, as no existing SFM has model weights, code, and training data publicly available under open-source terms. In this work, we take the first step toward filling this gap by focusing on the 24 official languages of the European Union (EU). We collect suitable training data by surveying automatic speech recognition datasets and unlabeled speech corpora under open-source compliant licenses, for a total of 950k hours. Additionally, we release automatic transcripts for 441k hours of unlabeled data under the permissive CC-BY license, thereby facilitating the creation of open-source SFMs for the EU languages.

语音模型开源数据欧盟语言基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。