用知识蒸馏提升尼日利亚语语音识别,显著降低错误率。
Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation
- 两阶段蒸馏:先用单语模型引导,再用伪标签迭代优化。
- 平均相对词错误率下降29%,超越现有主流多语言模型。
- 开源5种尼日利亚语言模型,适合非洲语言研究者使用。
尽管现代多语言自动语音识别(ASR)系统支持多种尼日利亚语言,但其性能始终落后于英语、法语等高资源语言。尼日利亚语言面临数据稀缺、拼写不一致、声调符号、方言差异、频繁语码转换及本地专有名词等建模挑战。为此,我们提出一种两阶段知识蒸馏的多语言ASR框架:首先,基于语言特异性N-gram语言模型,从现有单语模型进行学生-教师知识蒸馏;其次,利用伪标签数据进行迭代自提升以进一步提高精度。该方法显著缩小性能差距,在通用语音和Fleurs基准上均优于当前最优多语言模型,平均相对词错误率(WER)降低29%。我们发布Sometin Beta Pass Notin(SBPN)基础多语言模型,覆盖约鲁巴语、豪萨语、伊博语、尼日利亚皮钦语和尼日利亚英语,提供两种规模版本:SBPN-Base(120M参数)与SBPN-Large(600M参数)。通过开放发布这些基础模型,旨在为该地区丰富的语音与文化研究提供可用资源。
原文摘要 · Abstract (English)
Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind high-resource languages like English and French. Nigerian languages present unique modelling hurdles, including acute data scarcity, inconsistent orthography, tonal diacritics, diverse accents, frequent code-switching, and localized named entities. To address these challenges, we developed a multilingual ASR framework utilizing a two-stage distillation process. First, we employ student-teacher knowledge distillation from existing monolingual models, conditioned on robust language-specific N-gram language models. Second, we perform iterative self improvement using pseudo-labelled data to further refine accuracy. Our method significantly bridges the performance gap, achieving on average a relative Word Error Rate (WER) reduction of 29 % over monolingual baselines. Our models also outperform state-of-the-art multilingual models across major benchmarks, including Common Voice and Fleurs. We introduce Sometin Beta Pass Notin (SBPN), a foundational multilingual ASR model covering Yorùbá, Hausa, Igbo, Nigerian Pidgin, and Nigerian English. SBPN is released in two sizes: SBPN-Base (120 M parameters) and SBPN-Large (600 M parameters). By releasing these as open foundation models, we aim to provide ASR resources for further research into the rich phonetic and cultural landscape of the region.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。