arXiv:2608.16379cs.CLcs.SD2026-08

跨书写系统评估库尔德语语音识别,发现直接评分会夸大错误率。

Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis

  • 统一参考文本的分段归一化设计,避免书写系统差异干扰评估
  • 阿拉伯文输出转拉丁文后误差率下降13.85点,字符错误率降49.72点
  • 残余错误多来自评分管道限制,非模型本身缺陷,适合语音评估研究者

在使用拉丁字母拼写的格鲁西库尔德语数据集上评估基于阿拉伯字母输出的语音识别模型,面临测量难题:直接评分将书写系统差异误判为识别错误。通过统一参考与假设的归一化处理可避免此问题,但会改变参考文本分词,混淆准确率提升与评分基数变化。本文在5名说话人共1,722个问卷片段(9,763个参考词元,117.9分钟)上,对未微调的MMS-1B-all-ckb模型进行测试。采用统一参考设计:参考文本固定为9,763词元,仅假设表示变化。原始阿拉伯文输出得WER 111.70%、CER 100.92%,无完全匹配词;拉丁转写后降至WER 102.36%、CER 57.89%;归一化至参考文本格式后进一步降至WER 97.85%、CER 51.20%。RAW→FOLDED使误差率下降13.85点(WER)和49.72点(CER),其中归一化贡献4.51点(WER)和6.69点(CER)。仍存在显著误差:仅14.53%参考词元完全匹配,编辑以替换为主,短片段的段级WER更高。南库尔德语微调系统(aranemini/southern-kurdish-asr)在相同设计下表现更差(1,703段,WER 109.56%,CER 55.85%),且有12,330个输出字符超出归一化表,需重新计算。MMS输出含613个无法转换或映射的字符,表明部分残余误差源于评分流程而非识别能力。本文将发布固定参考文本及逐段结果,受源语料共享条款约束。

原文摘要 · Abstract (English)

Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measurement problem before a modelling one: direct scoring treats writing-system differences as recognition errors. Jointly normalizing reference and hypothesis avoids this, but also changes reference tokenization, mixing agreement gains with a change in the scoring denominator. I evaluate MMS-1B-all with the Central Kurdish (ckb) adapter, used as released without adaptation, on 1,722 Garrusi questionnaire segments from five speakers (9,763 reference word tokens; 117.9 minutes). I use a common-reference design: the reference is folded once and fixed at 9,763 tokens, while only the hypothesis representation varies. The raw Arabic-script hypothesis scores 111.70% WER and 100.92% CER, with zero exact word matches. Latin transliteration gives 102.36% WER and 57.89% CER; folding it into the reference's reduced orthography gives 97.85% and 51.20%. Thus RAW-to-FOLDED reduces measured WER by 13.85 points and CER by 49.72 points; folding alone accounts for 4.51 and 6.69 points. Substantial error remains: 14.53% of reference tokens are exact matches, edits are substitution-dominated, and per-segment WER is higher for shorter segments. A Southern Kurdish fine-tuned system (aranemini/southern-kurdish-asr), scored under the same design, performs worse on every speaker (1,703 segments), with 109.56% WER and 55.85% CER. However, 12,330 output characters fall outside the folding table, so these rates must be recomputed against the corrected fixed reference. The MMS output also contains 613 unconverted or unmapped characters, showing that part of the residual error reflects scoring-pipeline limits rather than recognition alone. I will release the fixed reference and segment-level results, subject to source-corpus sharing terms, to support independent checking.

语音识别多语言评估方法库尔德语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。