解决非拉丁语系语音识别错误分析中的对齐难题
Breaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin Scripts
- 提出跨语言、无脚本依赖的自动对齐方法
- 在三种书写系统中实现一致的词性级错误分析
- 可直接用于提升语音识别模型性能
语音识别(ASR)系统常使用词错误率(WER)等综合指标评估,但无法捕捉错误的语法结构。细粒度分析如按词性(PoS)分类错误,需准确对齐识别结果与参考转录。然而现有对齐工具在非拉丁语系语言中表现不可靠。本文提出一种鲁棒、自动、语言无关的对齐机制,适用于不同ASR架构及拉丁与非拉丁书写系统。该方法实现了假设、参考与评估序列的一致对齐,为下游语言学分析奠定基础。在此基础上,我们使用标准词性标注器实现可扩展、可复现的词性级错误分析。特别地,我们在三种主要分词书写系统中开展分析:元音附标文字(泰米尔语、印地语、卡纳达语)、字母文字(英语、俄语、希腊语)和辅音音节文字(阿拉伯语)。此外,我们展示了如何利用此类错误信息改进训练过程,从而降低WER。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) systems are commonly evaluated using aggregate metrics such as Word Error Rate (WER), which do not capture the linguistic structure of errors. Fine-grained analysis, such as Part-of-Speech (PoS)-wise error characterization, requires accurate alignment between ASR hypotheses and reference transcriptions. However, existing alignment tools are often unreliable for languages written in non-Latin scripts. In this work, we address this gap by proposing a robust, automated, language-agnostic alignment mechanism applicable across ASR architectures and across languages written in both Latin and non-Latin scripts. This enables consistent alignment of hypotheses, references, and evaluation sequences, forming the basis for downstream linguistic analysis. Building on this, we employ standard PoS taggers to perform scalable and reproducible PoS-wise error analysis. Notably, we perform alignment and downstream ASR error analysis across three major segmented writing systems, namely, Abugida (Tamil, Hindi, Kannada), Alphabetic (English, Russian, Greek), and Abjad (Arabic). We further demonstrate how such error information can be leveraged during ASR training to improve metrics such as WER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。