通过句级方法识别儿童语音的可靠识别结果,提升应用准确性
Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
- 基于句级特征区分可靠与不可靠语音转录结果
- 在英荷语数据上实现超97%精准率,可筛选21%-56%低错误率语音
- 适合需要高可靠语音识别的应用场景如语言学习系统
自动语音识别(ASR)在儿童语言学习和读写能力培养等应用中日益普及,但高错误率限制了其效果。通过提前识别可靠的ASR输出可缓解负面影响。本文提出两种新的句级方法,分别针对朗读语音和对话语音,以筛选出可靠转录结果。在英语和荷兰语数据集上,使用基线和微调模型进行评估,结果显示,最佳策略对朗读和对话语音均能达到超过97.4%的精确率;采用当前最优策略,可自动筛选出21.0%至55.9%的语音数据,且这些数据的未识别错误率(UER)低于2.6%。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) is increasingly used in applications involving child speech, such as language learning and literacy acquisition. However, the effectiveness of such applications is limited by high ASR error rates. The negative effects can be mitigated by identifying in advance which ASR-outputs are reliable. This work aims to develop two novel approaches for selecting reliable ASR-output at the utterance level, one for selecting reliable read speech and one for dialogue speech material. Evaluations were done on an English and a Dutch dataset, each with a baseline and finetuned model. The results show that utterance-level selection methods for identifying reliably transcribed speech recordings have high precision for the best strategy (P > 97.4) for both read speech and dialogue material, for both languages. Using the current optimal strategy allows 21.0% to 55.9% of dialogue/read speech datasets to be automatically selected with low (UER of < 2.6) error rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。