用海量儿童语音训练出跨语言语音识别模型,助力早期语言发育研究
BabAR: from phoneme recognition to developmental measures of young children's speech production
- 基于多语言儿童语音预训练,结合20秒上下文提升识别准确率
- 在五种语言上实现高精度语音识别,错误多属同类别替换
- 适合研究儿童语言发育、语音分析与早期干预的学者使用
大规模研究早期语言发育需要自动化工具,但针对幼儿的自动音素识别仍面临挑战。基于数十年数据积累,我们构建了TinyVox——一个包含超过50万条英语、法语、葡萄牙语、德语和西班牙语儿童语音片段的音素标注语料库。利用TinyVox训练了BabAR,一种面向儿童语音的跨语言音素识别系统。实验表明,在多语言儿童主导的全天录音上进行预训练显著优于其他方法,且在微调时提供20秒音频上下文可进一步提升性能。错误分析显示,错误主要发生在同一类音素之间,表明该系统适用于粗粒度的语言发展评估。通过验证,BabAR生成的语音成熟度指标与文献中的发育估计高度一致。
原文摘要 · Abstract (English)
Studying early speech development at scale requires automatic tools, yet automatic phoneme recognition, especially for young children, remains largely unsolved. Building on decades of data collection, we curate TinyVox, a corpus of more than half a million phonetically transcribed child vocalizations in English, French, Portuguese, German, and Spanish. We use TinyVox to train BabAR, a cross-linguistic phoneme recognition system for child speech. We find that pretraining the system on multilingual child-centered daylong recordings substantially outperforms alternatives, and that providing 20 seconds of surrounding audio context during fine-tuning further improves performance. Error analyses show that substitutions predominantly fall within the same broad phonetic categories, suggesting suitability for coarse-grained developmental analyses. We validate BabAR by showing that its automatic measures of speech maturity align with developmental estimates from the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。