arXiv:2602.04716cs.CL2026-02被引 2

用音位特征误差揭示非洲语言语音识别的隐藏问题

Linguistically Informed Evaluation of Multilingual ASR for African Languages

  • 引入音位特征误差(FER)和声调误差(TER)替代传统词错误率
  • 发现模型在音段特征上表现较好,但声调(尤其中调与降阶)仍难捕捉
  • 适合关注非洲语言语音识别、评估方法改进的研究者

词错误率(WER)将非洲语言的音韵、声调及其他语言学错误合并为单一词汇错误,导致性能误判。相比之下,特征错误率(FER)能揭示模型中具有语言学意义的错误。本文通过在两种非洲语言上补充字符错误率(CER)和FER,并引入声调感知扩展(TER),评估三种语音编码器。结果显示,即使词级准确率低,基于音位特征的FER与TER仍能揭示显著的语言学误差模式。模型在音段特征上表现更优,而声调(特别是中调与降阶)仍是最大挑战。以约鲁巴语为例,其WER=0.788,CER=0.305,FER=0.151;对于未参与预训练的濒危语言恩梅语,模型虽近乎全错(WER接近1),但CER=0.461,FER仅为0.267,表明错误多源于单个音素特征失误,被传统全有或全无指标掩盖。

原文摘要 · Abstract (English)

Word Error Rate (WER) mischaracterizes ASR models' performance for African languages by combining phonological, tone, and other linguistic errors into a single lexical error. By contrast, Feature Error Rate (FER) has recently attracted attention as a viable metric that reveals linguistically meaningful errors in models' performance. In this paper, we evaluate three speech encoders on two African languages by complementing WER with CER, and FER, and add a tone-aware extension (TER). We show that by computing errors on phonological features, FER and TER reveal linguistically-salient error patterns even when word-level accuracy remains low. Our results reveal that models perform better on segmental features, while tones (especially mid and downstep) remain the most challenging features. Results on Yoruba show a striking differential in metrics, with WER=0.788, CER=0.305, and FER=0.151. Similarly for Uneme (an endangered language absent from pretraining data) a model with near-total WER and 0.461 CER achieves the relatively low FER of 0.267. This indicates model error is often attributable to individual phonetic feature errors, which is obscured by all-or-nothing metrics like WER.

语音识别非洲语言评估指标音位特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。