arXiv:2506.01439cs.CLeess.AS2025-06被引 2

Whale模型融合w2v-BERT与E-Branchformer,在多语种语音识别上达到领先水平。

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

  • 采用w2v-BERT预训练+E-Branchformer编码器,结合CTC-注意力联合解码。
  • 在Librispeech测试集上词错误率低至2.4%,在CSJ数据集上字符错误率3.4%。
  • 适合需要高鲁棒性的多语种语音识别场景,如跨语言语音助手。

本文介绍了大规模语音识别模型Whale的开发。与Whisper和OWSM等模型类似,Whale兼具大模型规模和多样化的海量数据。其架构融合了w2v-BERT自监督模型、基于E-Branchformer的编码器-解码器主干网络,以及联合的CTC-注意力解码策略。训练数据涵盖公开数据集和内部数据,提升了对不同发音风格和声学环境的鲁棒性。在多个基准测试中,Whale表现优异,尤其在Librispeech test-clean集上实现2.4%的词错误率,在CSJ eval3集上达到3.4%的字符错误率,优于Whisper large-v3和OWSM v3.1。

原文摘要 · Abstract (English)

This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a diverse, extensive dataset. Whale's architecture integrates w2v-BERT self-supervised model, an encoder-decoder backbone built on E-Branchformer, and a joint CTC-attention decoding strategy. The training corpus comprises varied speech data, of not only public corpora but also in-house data, thereby enhancing the model's robustness to different speaking styles and acoustic conditions. Through evaluations on multiple benchmarks, Whale achieved comparable performance to existing models. In particular, it achieves a word error rate of 2.4% on the Librispeech test-clean set and a character error rate of 3.4% on the CSJ eval3 set, outperforming Whisper large-v3 and OWSM v3.1.

语音识别多语言自监督E-Branchformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。