arXiv:2608.10670cs.CL2026-08中稿 · oral presentation …

低资源方言语音识别需多种子评估,避免偶然误差。

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

论文配图:Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR
图 1 · 摘自论文原文
  • 在官方数据集上用多种子训练,验证结果可靠性
  • 标准CTC配合w2v-BERT 2.0达47.0% WER,优于大模型
  • 适合关注低资源语音识别可复现性的研究者

针对低资源方言典型语料规模下单次实验结果不可复现的问题,本文以印度喜马拉雅山区的加尔瓦利语为例,构建首个基于官方VAANI划分的可复现多种子语音识别基准,提供每种子的输出结果并进行显著性检验。重新评估潜在提升时发现:焦点CTC和音节加权目标均未在种子级别测试中优于标准CTC,音节目标甚至未能减少其针对性错误;从印地语迁移学习也未带来优于直接微调的收益。真正稳健的是基础方案:使用w2v-BERT 2.0与标准CTC,在五个种子上平均达到47.0%的词错误率(WER),优于更大的MMS-1B及同类模型;性能关键在于预训练设计而非参数量,速度增强带来小而一致的改善。多种子评估有效区分真实性能提升与随机噪声。

原文摘要 · Abstract (English)

At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first reproducible multi-seed ASR benchmark on the official VAANI splits, with per-seed outputs and significance testing. Re-examining plausible gains, we find them fragile: neither Focal CTC nor a matra-weighted objective beats standard CTC under seed-level testing, the matra objective fails to cut even its targeted errors, and Hindi-to-Garhwali transfer gives no gain over direct fine-tuning. What holds up is mundane: w2v-BERT 2.0 with standard CTC reaches 47.0% WER over five seeds, beating the larger MMS-1B and comparable models; pretraining design, not parameter count, drives performance, and speed augmentation gives a small, largely consistent gain. Multi-seed evaluation on official splits separates real gains from seed noise.

语音识别低资源语言可复现性多种子评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。