早层解码能提升萨丁尼亚语语音识别准确率
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
- 用逐层解码分析中间层比最终层更准
- 去掉顶层后字错误率下降,最优在倒数第二层
- 适合低资源语言研究,可发现表面指标忽略的问题
近期研究表明,多语言语音模型的中间层常比最终输出层包含更准确的音素表征。本文针对预训练Wav2Vec2模型,采用逐层解码策略,研究音素级预测在编码器各层间的演化过程,聚焦低资源语言卡米丹内塞萨丁尼亚语。结果表明,截断顶层可降低音素错误率(PER),最佳性能出现在倒数第二层而非最终层。通过精细对齐分析发现,中间层预测更保留学段身份,避免过生成,并减少特定类型音系错误。我们提出‘逆向错误’概念——指中间层正确预测被最终层错误覆盖的情况。这类现象揭示了表层错误指标的局限性,说明深层可能过度泛化而丢失声学细节。研究支持在低资源场景中使用早层探测作为诊断工具,以捕捉标准评估无法反映的语言学行为。
原文摘要 · Abstract (English)
Recent studies have shown that intermediate layers in multilingual speech models often encode more phonetically accurate representations than the final output layer. In this work, we apply a layer-wise decoding strategy to a pretrained Wav2Vec2 model to investigate how phoneme-level predictions evolve across encoder layers, focusing on Campidanese Sardinian, a low-resource language. We show that truncating upper transformer layers leads to improved Phoneme Error Rates (PER), with the best performance achieved not at the final layer, but two layers earlier. Through fine-grained alignment analysis, we find that intermediate predictions better preserve segmental identity, avoid overgeneration, and reduce certain classes of phonological errors. We also introduce the notion of regressive errors, cases where correct predictions at intermediate layers are overwritten by errors at the final layer. These regressions highlight the limitations of surface-level error metrics and reveal how deeper layers may generalize or abstract away from acoustic detail. Our findings support the use of early-layer probing as a diagnostic tool for ASR models, particularly in low-resource settings where standard evaluation metrics may fail to capture linguistically meaningful behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。