arXiv:2506.01192eess.AScs.SD2025-06中稿 · Interspeech 2025被引 5

用语音识别模型生成掩码目标,实现高效自监督语音识别

GigaAM: Efficient Self-Supervised Learner for Speech Recognition

  • 用语音识别模型生成掩码目标,提升自监督学习效果
  • 训练出的GigaAM模型在俄语识别上比Whisper-large-v3高50%
  • 支持全上下文与流式推理,适合工业级语音系统

自监督学习(SSL)在语音处理中表现出色,尤其在自动语音识别(ASR)领域。本文提出一种基于语音识别模型生成掩码目标的自监督预训练框架,并引入分块注意力机制,结合动态分块大小采样,实现预训练阶段同时支持全上下文与流式微调。实验探究了模型规模与数据量的扩展性。基于该方法,我们训练了GigaAM系列模型,其中一项在俄语语音识别上达到新纪录,性能优于Whisper-large-v3达50%。相关基础模型、语音识别模型及推理代码已开源,许可为MIT,可在https://github.com/salute-developers/gigaam获取。

原文摘要 · Abstract (English)

Self-Supervised Learning (SSL) has demonstrated strong performance in speech processing, particularly in automatic speech recognition. In this paper, we explore an SSL pretraining framework that leverages masked language modeling with targets derived from a speech recognition model. We also present chunkwise attention with dynamic chunk size sampling during pretraining to enable both full-context and streaming fine-tuning. Our experiments examine scaling with respect to model size and the amount of data. Using our method, we train the GigaAM family of models, including a state-of-the-art model for Russian speech recognition that outperforms Whisper-large-v3 by 50%. We have released our foundation and ASR models, along with the inference code, under the MIT license as open-source resources to the research community. Available at https://github.com/salute-developers/gigaam.

自监督学习语音识别大模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。