通过因果发现解析儿童语音识别错误的多重影响因素
Causal Structure Discovery for Error Diagnostics of Children's ASR
- 构建因果结构模型揭示生理、认知、外部因素间的相互作用
- 量化显示年龄对识别错误有直接与间接双重影响
- 适用于不同语音模型,适合语音识别优化研究者
儿童自动语音识别(ASR)性能常低于成人,受生理(如声道较小)、认知(如发音不成熟)和外部因素(如词汇量小、背景噪声)等多重因素交织影响。现有分析方法孤立考察各因素,忽视其相互依赖关系,例如年龄既直接影响识别准确率,又通过发音能力间接影响。本文提出因果结构发现方法,揭示生理、认知、外部因素与ASR错误间的复杂关联;进一步采用因果量化评估各因素影响程度,并扩展至微调模型,识别哪些因素经微调后改善、哪些仍持续存在。在Whisper和Wav2Vec2.0上的实验表明,该方法可泛化至不同ASR系统。
原文摘要 · Abstract (English)
Children's automatic speech recognition (ASR) often underperforms compared to that of adults due to a confluence of interdependent factors: physiological (e.g., smaller vocal tracts), cognitive (e.g., underdeveloped pronunciation), and extrinsic (e.g., vocabulary limitations, background noise). Existing analysis methods examine the impact of these factors in isolation, neglecting interdependencies-such as age affecting ASR accuracy both directly and indirectly via pronunciation skills. In this paper, we introduce a causal structure discovery to unravel these interdependent relationships among physiology, cognition, extrinsic factors, and ASR errors. Then, we employ causal quantification to measure each factor's impact on children's ASR. We extend the analysis to fine-tuned models to identify which factors are mitigated by fine-tuning and which remain largely unaffected. Experiments on Whisper and Wav2Vec2.0 demonstrate the generalizability of our findings across different ASR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。