首个阿拉伯语儿童语音数据集,揭示现有模型在儿童语音上表现严重下降
Arabic Little STT: Arabic Children Speech Recognition Dataset
- 构建了355条阿拉伯语儿童语音,来自288名6-13岁学生
- 最强模型在儿童语音上仍达0.66的词错误率,远高于成人数据集
- 呼吁建立儿童专用语音基准,推动语音技术公平性
人工智能系统性能高度依赖高质量训练数据,但阿拉伯语等低资源语言面临严重数据匮乏。尤其缺乏专为儿童设计的语音语料库,造成显著挑战。为此,我们构建了阿拉伯语儿童语音数据集 Arabic Little STT,包含355条来自288名6至13岁儿童的课堂录音。我们对当前最先进的自动语音识别(ASR)模型 Whisper 在该数据集上的表现进行了系统评估,并与成人阿拉伯语基准对比。八种 Whisper 变体结果显示,即使最佳模型(Large_v3)在儿童语音上词错误率(WER)仍高达0.66,远高于其在成人数据集上低于0.20的表现。该结果与英语研究一致,凸显了为儿童语音建立专用基准的迫切需求。同时强调,此类数据必须在严格伦理与隐私框架下管理,以保护儿童敏感信息。本研究旨在为阿拉伯语儿童语音技术的公平发展迈出第一步,公开数据集有助于提升儿童在语音识别数据中的代表性。
原文摘要 · Abstract (English)
The performance of Artificial Intelligence (AI) systems fundamentally depends on high-quality training data. However, low-resource languages like Arabic suffer from severe data scarcity. Moreover, the absence of child-specific speech corpora is an essential gap that poses significant challenges. To address this gap, we present our created dataset, Arabic Little STT, a dataset of Levantine Arabic child speech recorded in classrooms, containing 355 utterances from 288 children (ages 6 - 13). We further conduct a systematic assessment of Whisper, a state-of-the-art automatic speech recognition (ASR) model, on this dataset and compare its performance with adult Arabic benchmarks. Our evaluation across eight Whisper variants reveals that even the best-performing model (Large_v3) struggles significantly, achieving a 0.66 word error rate (WER) on child speech, starkly contrasting with its sub 0.20 WER on adult datasets. These results align with other research on English speech. Results highlight the critical need for dedicated child speech benchmarks and inclusive training data in ASR development. Emphasizing that such data must be governed by strict ethical and privacy frameworks to protect sensitive child information. We hope that this study provides an initial step for future work on equitable speech technologies for Arabic-speaking children. We hope that our publicly available dataset enrich the children's demographic representation in ASR datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。