评测语音识别系统需结合多种指标,避免仅看平均错误率。
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
- 对比多种性能与偏见度量方法,评估荷兰语端到端语音识别系统。
- 发现平均错误率无法全面反映系统对不同人群的识别差异。
- 建议报告多维度指标,提升对多样性群体的表现评估透明度。
越来越多证据表明,自动语音识别(ASR)系统对不同说话人及群体存在偏差,如性别、年龄或口音差异。当前研究主要聚焦于检测与量化偏差,并开发缓解方法。然而,如何有效衡量系统性能与偏见仍是开放问题。本研究对比了文献中及新提出的多种性能与偏见度量方法,评估针对荷兰语的先进端到端ASR系统。实验采用多种偏见缓解策略处理不同说话人群体的偏差。结果表明,仅依赖平均错误率不足以全面评估系统表现,需辅以其他指标。论文最后提出报告ASR性能与偏见的改进建议,以更真实反映系统在多样化说话人中的表现及整体偏见水平。
原文摘要 · Abstract (English)
There is increasingly more evidence that automatic speech recognition (ASR) systems are biased against different speakers and speaker groups, e.g., due to gender, age, or accent. Research on bias in ASR has so far primarily focused on detecting and quantifying bias, and developing mitigation approaches. Despite this progress, the open question is how to measure the performance and bias of a system. In this study, we compare different performance and bias measures, from literature and proposed, to evaluate state-of-the-art end-to-end ASR systems for Dutch. Our experiments use several bias mitigation strategies to address bias against different speaker groups. The findings reveal that averaged error rates, a standard in ASR research, alone is not sufficient and should be supplemented by other measures. The paper ends with recommendations for reporting ASR performance and bias to better represent a system's performance for diverse speaker groups, and overall system bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。