arXiv:2501.06478eess.AScs.CL2025-01中稿 · ICASSP 2025被引 4

用少量儿童语音数据,实现对南非语和科萨语幼儿口语故事的自动识别。

Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives

  • 基于Whisper模型,仅需5分钟儿童语音即可训练。
  • 加入同主题成人语音数据+语音转换,性能提升最明显。
  • 适合研究非洲语言儿童语言发育或低资源语音识别者。

我们开发了针对阿非利卡语和科萨语学龄前儿童口语叙事的自动语音识别(ASR)系统。口头叙事可在儿童识字前评估其语言发展水平。本研究考察了多种先前的儿童语音ASR策略,以确定最适合该场景的方法。实验发现,仅使用5分钟领域内儿童语音标注数据,结合同领域成人语音数据与语音转换技术,可带来最大性能提升;半监督学习对两种语言均有帮助,而参数高效微调仅在阿非利卡语中有效(科萨语在Whisper模型中代表性不足)。现有儿童语音研究多聚焦英语,且极少关注4至5岁幼儿。本工作首次在非英语、低资源语言及学前阶段验证了多种儿童语音识别策略的有效性。

原文摘要 · Abstract (English)

We develop automatic speech recognition (ASR) systems for stories told by Afrikaans and isiXhosa preschool children. Oral narratives provide a way to assess children's language development before they learn to read. We consider a range of prior child-speech ASR strategies to determine which is best suited to this unique setting. Using Whisper and only 5 minutes of transcribed in-domain child speech, we find that additional in-domain adult data (adult speech matching the story domain) provides the biggest improvement, especially when coupled with voice conversion. Semi-supervised learning also helps for both languages, while parameter-efficient fine-tuning helps on Afrikaans but not on isiXhosa (which is under-represented in the Whisper model). Few child-speech studies look at non-English data, and even fewer at the preschool ages of 4 and 5. Our work therefore represents a unique validation of a wide range of previous child-speech ASR strategies in an under-explored setting.

语音识别儿童语音低资源语言语言发育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。