arXiv:2505.22759cs.CLcs.AI2025-05被引 2

首个面向英意语的开源语音基础模型,训练数据超150小时。

FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian

  • 基于150万+小时开源语音数据训练,完全公开代码与数据。
  • 性能媲美现有模型,推理速度最高快8倍。
  • 适合关注语音技术透明性与可复现性的研究者。

语音基础模型(SFMs)如Whisper和SeamlessM4T显著推动了语音处理的发展。然而,其封闭特性——训练数据与代码不可获取——带来了严重的可复现性和公平评估挑战。尽管其他领域已通过开源代码和数据实现全面透明,但语音领域的类似进展仍有限。为此,我们提出FAMA,首个面向英语和意大利语的开源科学语音基础模型家族,训练数据超过150,000小时。此外,我们还发布了一个包含16,000小时清洗并伪标注语音的新数据集。结果表明,FAMA在性能上达到现有模型水平,同时推理速度最高提升8倍。所有成果,包括代码、数据集和模型,均以符合开源协议的方式发布,推动语音技术研究的开放化进程。

原文摘要 · Abstract (English)

The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inaccessible training data and code--poses major reproducibility and fair evaluation challenges. While other domains have made substantial progress toward open science by developing fully transparent models trained on open-source (OS) code and data, similar efforts in speech remain limited. To fill this gap, we introduce FAMA, the first family of open science SFMs for English and Italian, trained on 150k+ hours of OS speech data. Moreover, we present a new dataset containing 16k hours of cleaned and pseudo-labeled speech for both languages. Results show that FAMA achieves competitive performance compared to existing SFMs while being up to 8 times faster. All artifacts, including code, datasets, and models, are released under OS-compliant licenses, promoting openness in speech technology research.

语音模型开源基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。