arXiv:2506.08999cs.CLcs.AI2025-06被引 6

用自监督学习提升跨语言儿童语音成熟度识别能力

Employing self-supervised learning models for cross-linguistic child speech maturity classification

  • 构建包含24万条标注语音的跨语言数据集SpeechMaturity
  • 模型准确率接近人类水平,且在城乡场景中均表现稳定
  • 适合从事儿童语音分析与多语言语音识别的研究者

语音技术系统在儿童语音的下游任务中面临挑战,主要由于训练数据少且儿童语音本身复杂。本文引入新型数据集SpeechMaturity,应用于先进Transformer模型,解决一项基础分类任务:识别儿童发声类型。该数据集涵盖美国、玻利维亚、瓦努阿图、巴布亚新几内亚、所罗门群岛和法国等地区超过25种语言的儿童语音,样本量空前,共包含242,004条标注语音,远超以往工作。模型需区分哭声、笑声、成熟语音(辅音+元音)及不成熟语音(仅辅音或元音)。在该数据集上训练的模型优于此前最先进模型,在城乡环境中均保持鲁棒性,分类准确率接近人类水平。

原文摘要 · Abstract (English)

Speech technology systems struggle with many downstream tasks for child speech due to small training corpora and the difficulties that child speech pose. We apply a novel dataset, SpeechMaturity, to state-of-the-art transformer models to address a fundamental classification task: identifying child vocalizations. Unlike previous corpora, our dataset captures maximally ecologically-valid child vocalizations across an unprecedented sample, comprising children acquiring 25+ languages in the U.S., Bolivia, Vanuatu, Papua New Guinea, Solomon Islands, and France. The dataset contains 242,004 labeled vocalizations, magnitudes larger than previous work. Models were trained to distinguish between cry, laughter, mature (consonant+vowel), and immature speech (just consonant or vowel). Models trained on the dataset outperform state-of-the-art models trained on previous datasets, achieved classification accuracy comparable to humans, and were robust across rural and urban settings.

儿童语音自监督学习跨语言语音分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。