提出新评估指标PIER,更精准衡量代码转换语音识别效果。
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
- 用特定关注词构建新型误差率指标PIER,聚焦关键词汇
- 传统指标在代码转换上表现虚假提升,实际效果不佳
- 适合研究代码转换语音识别的学者与开发者参考
代码转换(语言在单个话语中交替)对自动语音识别构成重大挑战。尽管任务特殊,性能仍常以通用指标如词错误率(WER)衡量。本文通过连接时序分类和编码器-解码器模型发现:在非代码转换数据上微调后,经典指标在代码转换测试集上表现提升,但实际代码转换词错误率反而上升(符合预期)。为此,我们提出关注点误差率(PIER),一种仅针对特定关注词计算的WER变体。在代码转换语句上实例化PIER后,其能更准确反映模型真实性能,揭示未来工作巨大改进空间。该方法可更精确评估模型在跨词与词内代码转换等复杂场景的表现。
原文摘要 · Abstract (English)
Code-switching, the alternation of languages within a single discourse, presents a significant challenge for Automatic Speech Recognition. Despite the unique nature of the task, performance is commonly measured with established metrics such as Word-Error-Rate (WER). However, in this paper, we question whether these general metrics accurately assess performance on code-switching. Specifically, using both Connectionist-Temporal-Classification and Encoder-Decoder models, we show fine-tuning on non-code-switched data from both matrix and embedded language improves classical metrics on code-switching test sets, although actual code-switched words worsen (as expected). Therefore, we propose Point-of-Interest Error Rate (PIER), a variant of WER that focuses only on specific words of interest. We instantiate PIER on code-switched utterances and show that this more accurately describes the code-switching performance, showing huge room for improvement in future work. This focused evaluation allows for a more precise assessment of model performance, particularly in challenging aspects such as inter-word and intra-word code-switching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。