剖析儿童语音识别错误成因,揭示年龄与音频词数是主要影响因素
Causal Analysis of ASR Errors for Children: Quantifying the Impact of Physiological, Cognitive, and Extrinsic Factors
- 基于因果推断分析生理、认知与外部因素对儿童语音识别的影响
- 年龄和音频词数影响最大,背景噪声与发音能力次之
- 微调可降低生理认知影响,但词数敏感性仍存,适合儿童语音研究者
近年来,儿童自动语音识别(ASR)系统应用日益广泛,推动了针对儿童语音模型准确率提升的研究。当前方法多直接使用开源语音基础模型(SFM)或在儿童语音数据上进行微调,但这些模型在儿童语音上的词错误率(WER)普遍高于成人语音。然而,对性能下降原因的系统性分析仍显不足。本研究通过两阶段分析填补该空白:第一阶段在两个儿童语音语料库上对两种自监督模型(Wav2Vec2.0、Hubert)和两种弱监督模型(Whisper、MMS)按不同年龄段进行基准测试,建立因果分析数据基础;第二阶段运用因果推断,分析生理因素(年龄、性别)、认知因素(发音能力)及外部因素(词汇难度、背景噪声、音频词数)对模型准确率的影响。结果表明,生理因素(年龄)和特定外部因素(音频词数)影响最大,其次为背景噪声和发音能力。在儿童语音上微调模型可降低对生理与认知因素的敏感性,但对词数的敏感性依然存在。
原文摘要 · Abstract (English)
The increasing use of children's automatic speech recognition (ASR) systems has spurred research efforts to improve the accuracy of models designed for children's speech in recent years. The current approach utilizes either open-source speech foundation models (SFMs) directly or fine-tuning them with children's speech data. These SFMs, whether open-source or fine-tuned for children, often exhibit higher word error rates (WERs) compared to adult speech. However, there is a lack of systemic analysis of the cause of this degraded performance of SFMs. Understanding and addressing the reasons behind this performance disparity is crucial for improving the accuracy of SFMs for children's speech. Our study addresses this gap by investigating the causes of accuracy degradation and the primary contributors to WER in children's speech. In the first part of the study, we conduct a comprehensive benchmarking study on two self-supervised SFMs (Wav2Vec2.0 and Hubert) and two weakly supervised SFMs (Whisper and MMS) across various age groups on two children speech corpora, establishing the raw data for the causal inference analysis in the second part. In the second part of the study, we analyze the impact of physiological factors (age, gender), cognitive factors (pronunciation ability), and external factors (vocabulary difficulty, background noise, and word count) on SFM accuracy in children's speech using causal inference. The results indicate that physiology (age) and particular external factor (number of words in audio) have the highest impact on accuracy, followed by background noise and pronunciation ability. Fine-tuning SFMs on children's speech reduces sensitivity to physiological and cognitive factors, while sensitivity to the number of words in audio persists. Keywords: Children's ASR, Speech Foundational Models, Causal Inference, Physiology, Cognition, Pronunciation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。