评估语音大模型在自闭症诊断对话中的表现,发现儿童语音识别率下降显著。
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
- 对比多个语音大模型在亲子对话中的表现,聚焦儿童语音识别难点。
- 儿童语音识别错误率比成人高15-20%,模型泛化能力受限。
- 用LoRA微调Whisper-large,在资源有限下提升8%-13%识别准确率。
临床环境中对儿童与成人对话的可靠转录对于自闭症等发育障碍的诊断至关重要。近年来,深度学习和大规模标注数据推动了语音基础模型的发展,显著提升了自动语音识别(ASR)性能。然而,这些模型在亲子对话场景中的表现仍缺乏系统评估。本文针对自闭症诊断会话中儿童-成人互动数据集,全面评估了Whisper、Wav2Vec2、HuBERT和WavLM的表现。结果发现,在对话设置下,语音基础模型对儿童语音的识别错误率(WER)相比成人语音有15-20%的绝对增长。随后,我们在低资源条件下使用LoRA微调表现最佳的零样本模型Whisper-large,使儿童和成人语音的WER分别降低8%和13%。
原文摘要 · Abstract (English)
Reliable transcription of child-adult conversations in clinical settings is crucial for diagnosing developmental disorders like Autism. Recent advances in deep learning and availability of large scale transcribed data has led to development of speech foundation models that have shown dramatic improvements in ASR performance. However, their performance on conversational child-adult interactions remains underexplored. In this work, we provide a comprehensive evaluation of ASR performance on a dataset containing child-adult interactions from autism diagnostic sessions, using Whisper, Wav2Vec2, HuBERT, and WavLM. We find that speech foundation models show a noticeable performance drop (15-20% absolute WER) for child speech compared to adult speech in the conversational setting. Then, we fine-tune the best-performing zero-shot model (Whisper-large) using LoRA in a low-resource setting, yielding 8% and 13% absolute WER improvements for child and adult speech, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。