用真实口语错误数据测试语音识别模型,发现WhisperX在错误纠正上表现有限。
Evaluating ASR robustness to spontaneous speech errors: A study of WhisperX using a Speech Error Database
- 基于口语错误数据库构建标注体系,覆盖词级与音节级错误位置
- 在5300个错误样本上测试WhisperX,准确率未达理想水平
- 适合评估语音识别系统对自然口误的鲁棒性,研究者可参考此框架
西蒙弗雷泽大学口语错误数据库(SFUSED)是为语言学和心理语言学研究开发的公开数据集。本文展示其设计与标注如何用于测试语音识别模型。该数据库包含来自自然英语口语的系统化标注错误,每个错误均标记了预期发音与实际发音。标注体系涵盖多个分类维度,包括语言层级、上下文敏感性、发音退化、词语修正以及词级与音节级错误定位,对模型评估具有价值。为验证这些分类变量的有效性,本文在5,300个记录的词汇与音位错误上评估了WhisperX的转录准确率。分析表明,该数据库可作为诊断语音识别系统性能的有效工具。
原文摘要 · Abstract (English)
The Simon Fraser University Speech Error Database (SFUSED) is a public data collection developed for linguistic and psycholinguistic research. Here we demonstrate how its design and annotations can be used to test and evaluate speech recognition models. The database comprises systematically annotated speech errors from spontaneous English speech, with each error tagged for intended and actual error productions. The annotation schema incorporates multiple classificatory dimensions that are of some value to model assessment, including linguistic hierarchical level, contextual sensitivity, degraded words, word corrections, and both word-level and syllable-level error positioning. To assess the value of these classificatory variables, we evaluated the transcription accuracy of WhisperX across 5,300 documented word and phonological errors. This analysis demonstrates the atabase's effectiveness as a diagnostic tool for ASR system performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。