构建首个大规模多语言语音模型,揭示其性能随规模变化的规律。
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
- 基于150种语言36万小时数据,训练0.25B到18B参数模型
- 发现规模扩大可显著提升低资源语言性能,缓解技术偏见
- 开源模型与数据,支持未来语音技术研究
神经尺度定律为构建稳健的序列处理架构提供了重要参考。尽管在其他模态中已得到充分研究,语音领域的尺度规律仍相对未被深入探索。本文提出OWLS,一个公开可复现的多语言语音识别与翻译模型套件,涵盖0.25B至18B参数,其中18B版本为目前已知最大语音模型。该套件利用跨150种语言的最高达36万小时的公共语音数据,系统研究数据、模型和计算量对多语言语音任务性能的影响。通过OWLS,我们推导出神经尺度定律,证明性能可被可靠预测。关键发现是:规模扩展能有效提升低资源语言/方言的表现,有助于缓解偏见,提升语音技术的可及性。最后,我们展示如何利用OWLS开启新研究方向,揭示大规模语音模型中的涌现能力。模型检查点将发布于https://huggingface.co/collections/espnet/owls-scaling-laws-for-speech-recognition-and-translation-67ab7f991c194065f057ce8d,供后续研究使用。
原文摘要 · Abstract (English)
Neural scaling laws offer valuable insights for designing robust sequence processing architectures. While these laws have been extensively characterized in other modalities, their behavior in speech remains comparatively underexplored. In this work, we introduce OWLS, an open-access, reproducible suite of multilingual speech recognition and translation models spanning 0.25B to 18B parameters, with the 18B version being the largest speech model, to the best of our knowledge. OWLS leverages up to 360K hours of public speech data across 150 languages, enabling a systematic investigation into how data, model, and compute scaling each influence performance in multilingual speech tasks. We use OWLS to derive neural scaling laws, showing how final performance can be reliably predicted when scaling. One of our key findings is that scaling enhances performance on low-resource languages/dialects, helping to mitigate bias and improve the accessibility of speech technologies. Finally, we show how OWLS can be used to power new research directions by discovering emergent abilities in large-scale speech models. Model checkpoints will be released on https://huggingface.co/collections/espnet/owls-scaling-laws-for-speech-recognition-and-translation-67ab7f991c194065f057ce8d for future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。