arXiv:2409.18584cs.SDeess.AS2024-09被引 15

构建3-5岁儿童普通话语音数据集,助力儿童语音识别研究。

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5

  • 收集397名儿童41.25小时语音,人工精标并覆盖多地
  • 微调预训练模型使识别准确率显著提升
  • 支持语音识别与说话人验证,适合儿童语音研究者

自动语音识别(ASR)系统在Whisper、Conformer及Wav2vec 2.0、HuBERT等模型推动下取得显著进展。然而,针对3至5岁儿童的普通话语音识别仍面临发音、语调和语速差异大等挑战。本文提出一个面向该年龄段的中文语音数据集ChildMandarin,包含41.25小时经人工精标语音,来自中国多个省份的397名儿童,性别分布均衡。我们分析了说话人人口统计特征、语音时长分布与地理覆盖情况。实验表明,在从头训练的Conformer模型以及微调HuBERT和Whisper等预训练模型上,微调策略显著提升识别性能。此外,本数据集在说话人验证任务中也表现出良好效果,尽管幼儿声学特征独特。该数据集已开源,供学术研究免费使用(https://github.com/flageval-baai/ChildMandarin)。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research. The dataset is now open-source and freely available for all academic purposes on https://github.com/flageval-baai/ChildMandarin.

语音识别儿童语音中文数据集开源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。