arXiv:2605.13087cs.CLcs.AI2026-05中稿 · Interspeech 2026

针对印地语和马拉雅拉姆语语音识别,提出分层级评测基准并优化训练策略。

Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition

论文配图:Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition
图 1 · 摘自论文原文
  • 构建四层级语音评测集,涵盖录音室到噪声合成数据。
  • 早起大幅参数更新可降低12个绝对点的词错误率。
  • 适合低资源语言语音识别研究者与模型优化工程师。

将多语言语音识别模型Whisper微调用于低资源语言时,虽提升朗读语音性能,却损害自然口语表现。为诊断此矛盾,我们提出Vividh-ASR,一个针对印地语和马拉雅拉姆语的复杂度分层基准,覆盖四个层级:录音室、广播、自然口语及合成噪声。通过控制学习率时机与课程顺序的研究发现,早期进行大参数更新可使整体词错误率(WER)降低12个绝对点,而从难到易的课程设置对自然口语有额外增益。由此启发出反向多阶段微调(R-MFT)训练方法,使参数量仅244M的Whisper模型达到甚至超过传统769M模型的表现。通过核相似性(CKA)与奇异值分解(SVD)的表征分析显示,有效训练方案集中于解码器的适应,保留预训练编码器的声学结构。我们已开源该基准与模型。

原文摘要 · Abstract (English)

Fine-tuning multilingual ASR models like Whisper for low-resource languages often improves read speech but degrades spontaneous audio performance. To diagnose this mismatch, we introduce Vividh-ASR, a complexity-stratified benchmark for Hindi and Malayalam across four tiers: studio, broadcast, spontaneous, and synthetic noise. Through a controlled study of learning-rate timing and curriculum ordering, we find that early large parameter updates improve global WER by 12 absolute points, while a hard-to-easy curriculum adds gains for spontaneous speech. These findings motivate reverse multi-stage fine-tuning (R-MFT), a training recipe that enables a parameter-efficient 244M Whisper model to match or exceed conventionally fine-tuned 769M counterparts. Representational analysis via CKA and SVD reveals effective schedules concentrate adaptation in the decoder, preserving the pre-trained encoder's acoustic geometry. We release the benchmark and models.

语音识别低资源语言模型优化Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。