arXiv:2604.04598cs.CL2026-04被引 3

首个公开的普什图语多语言语音识别基准,揭示模型在脚本输出上的严重缺陷

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation

  • 首次在公共数据集上评测多模型零样本语音识别与跨域性能
  • 零样本最佳结果为39.7% WER,但多数模型生成阿拉伯文而非普什图文
  • 发现脚本错误掩盖真实性能,适合低资源语言研究者参考

普什图语由约6000万至8000万人使用,但缺乏公开的多语言自动语音识别(ASR)基准。本文报告首个可复现的多模型评估,涵盖零样本ASR、脚本级失败及微调模型的跨域表现。在零样本测试中,十种模型(含所有七种Whisper尺寸、MMS-1B、SeamlessM4T-v2-large和OmniASR-CTC-300M)在FLEURS普什图语测试集与过滤后的Common Voice~24子集上评估,零样本Whisper WER范围为90%至297%,中等模型在Common Voice~24上达461%且出现解码器循环。SeamlessM4T在Common Voice~24上实现39.7%最优零样本WER(截至投稿),MMS-1B在FLEURS上为43.8%。脚本审计显示,无一Whisper模型在超过0.8%的语音中输出普什图文,而MMS-1B、SeamlessM4T和OmniASR均超93%脚本保真度;仅看WER会掩盖此根本性失败。跨域评估中,五种微调模型在分布外数据上性能从14%降至32.5%–59%,仅一种增强模型在两数据集上保持35.1%且无退化。字符级错误分析表明,普什图语特有音素(卷舌音与边擦音)占显著误差比例。所有评估限于朗读语音。论文提出五大结构性障碍并建议五项优先研究方向。

原文摘要 · Abstract (English)

Pashto is spoken by approximately 60--80 million people but has no published benchmarks for multilingual automatic speech recognition (ASR) on any shared public test set. This paper reports the first reproducible multi-model evaluation on public Pashto data, covering zero-shot ASR, script-level failure, and cross-domain evaluation of fine-tuned models. For zero-shot ASR, ten models (all seven Whisper sizes, MMS-1B, SeamlessM4T-v2-large, and OmniASR-CTC-300M) are evaluated on the FLEURS Pashto test set and a filtered Common Voice~24 subset; zero-shot Whisper WER ranges from 90% to 297%, with the medium model collapsing to 461% on Common Voice~24 consistent with decoder looping. SeamlessM4T achieves 39.7% WER on Common Voice~24 (the best zero-shot result reported to date, as of submission); MMS-1B achieves 43.8% on FLEURS. For script failure, a language-identification audit shows that no Whisper model produces Pashto-script output in more than 0.8% of utterances, while MMS-1B, SeamlessM4T, and OmniASR each exceed 93% Pashto-script fidelity; WER alone does not reveal this failure, since a model generating Arabic-script output on Pashto audio has not achieved ASR in any interpretable sense. For cross-domain evaluation, five fine-tuned Pashto ASR models are evaluated on both test sets: published WER figures of 14% degrade to 32.5--59% on out-of-distribution sets, while one augmented model achieves 35.1% on both sets with zero cross-domain degradation. Character-class error stratification confirms that Pashto-unique phonemes (the retroflex series and lateral fricatives) account for disproportionate error mass. All evaluations cover read speech only. Five structural impediments to cumulative progress are identified and five ordered research priorities are argued.

语音识别低资源语言脚本错误跨域评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。