arXiv:2608.08235eess.AS2026-08被引 1

打造65种印地语系语言的开源语音识别模型,覆盖低资源方言。

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

论文配图:SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages
图 1 · 摘自论文原文
  • 用自监督预训练+图文对齐+端到端微调三阶段构建多语言模型
  • 在65种语言上实现最低词错误率,尤其提升低资源语言表现
  • 首个公开评估的低资源与部落语言语音识别系统,适合本土化应用

印度语言环境涵盖700多种语言和数千种方言,但绝大多数自动语音识别(ASR)系统仅支持其中少数。我们提出SraVaani-1.0,一个覆盖65种印度语言和方言的多语言ASR模型,其中许多此前无公开可用或竞争性ASR系统。该模型基于FastConformer架构,从零训练,分三阶段:第一阶段在VAANI语料库的31,255小时无标签语音上进行自监督预训练,使用对比学习目标;第二阶段引入音频-图像表示对齐,利用VAANI语料库中配对的图像与语音,通过视觉上下文与口语内容的关系,使语音编码器学习更丰富的语义表征,提升下游识别性能,尤其在低资源语言上;第三阶段,使用混合令牌与持续时间转换器(TDT)-CTC解码器,在24个公共数据集整理的31,263小时多语言印度语音上进行端到端微调。我们在八个基准上对比了SraVaani-1.0与三种先进多语言ASR系统,结果表明其在大量语言-数据集组合上达到最低词错误率(WER),且在高资源语言上仍具竞争力。更重要的是,它是唯一在VAANI基准上评估、可为多种低资源及部落印度语言提供转录能力的开源模型。

原文摘要 · Abstract (English)

India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction of this diversity. We present SraVaani-1.0, a multilingual ASR model covering 65 Indian languages and dialects, many of which currently have no publicly available or competing ASR system. SraVaani-1.0 is built on a FastConformer architecture and trained from scratch through a three stage the first stage, we perform self-supervised pretraining on 31,255 hours of unlabelled speech from the VAANI corpus using a contrastive learning objective. In the second stage, we introduce an audio-image representation alignment stage that leverages the paired images and speech available in the VAANI corpus. This multimodal alignment encourages the speech encoder to learn semantically richer representations by exploiting the relationship between visual context and spoken content, thereby improving downstream recognition, particularly for low resource the final stage, the aligned encoder is fine-tuned end-to-end using a Hybrid Token-and-Duration Transducer (TDT)-CTC decoder on 31,263 hours of labelled multilingual Indian speech compiled from 24 public datasets spanning 65 languages and dialects. We evaluate SraVaani-1.0 against three state-of-the-art multilingual ASR systems across eight benchmarks. SraVaani-1.0 achieves the lowest word error rate (WER) on a large number of language-dataset pairs while remaining competitive with the best-performing systems on high resource importantly, it is the only open-source evaluated model that provides transcription capability for multiple low-resource and tribal Indian languages, which are assessed exclusively on the VAANI benchmark.

语音识别多语言低资源开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。