arXiv:2603.20255cs.CLcs.HC2026-03

构建首个面向阿拉伯语儿童的语音分类数据集,助力低资源语言教育AI研究。

Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education

  • 提出分层音频分类框架,先分组后精分类
  • 覆盖141类,46397个3-12岁儿童语音样本
  • 适合研究阿拉伯语儿童语音识别与教育AI的学者

近年来,基于语音的AI教育应用受到广泛关注,尤其针对儿童。然而,受限于公开可用数据集的缺乏,尤其是阿拉伯语等低资源语言的儿童语音研究仍十分有限。本文提出Abjad-Kids,一个专为幼儿园及小学阶段教育设计的阿拉伯语语音数据集,聚焦字母、数字和颜色的基础学习。该数据集包含46,397个来自3至12岁儿童的音频样本,涵盖141个类别。所有样本均在受控条件下录制,确保时长、采样率和格式一致。针对阿拉伯语音素内部相似性高、每类样本量少的问题,提出基于CNN-LSTM的分层音频分类方法。该方法将字母识别分解为两阶段:初始分组分类模型,再对各组进行专用分类器处理。评估了基于语言学的静态分组与基于聚类的动态分组两种策略,结果显示静态分组表现更优。对比传统机器学习与深度学习方法,证实了结合数据增强的CNN-LSTM模型的有效性。尽管结果理想,多数实验仍存在过拟合问题,可能源于样本量不足,即使经过数据增强与正则化。未来工作将聚焦于收集更多数据以缓解此问题。Abjad-Kids将公开共享,期望其丰富儿童语音数据集中的阿拉伯语代表性,成为阿拉伯语儿童语音分类研究的重要资源。

原文摘要 · Abstract (English)

Speech-based AI educational applications have gained significant interest in recent years, particularly for children. However, children speech research remains limited due to the lack of publicly available datasets, especially for low-resource languages such as Arabic.This paper presents Abjad-Kids, an Arabic speech dataset designed for kindergarten and primary education, focusing on fundamental learning of alphabets, numbers, and colors. The dataset consists of 46397 audio samples collected from children aged 3 - 12 years, covering 141 classes. All samples were recorded under controlled specifications to ensure consistency in duration, sampling rate, and format. To address high intra-class similarity among Arabic phonemes and the limited samples per class, we propose a hierarchical audio classification based on CNN-LSTM architectures. Our proposed methodology decomposes alphabet recognition into a two-stage process: an initial grouping classification model followed by specialized classifiers for each group. Both strategies: static linguistic-based grouping and dynamic clustering-based grouping, were evaluated. Experimental results demonstrate that static linguistic-based grouping achieves superior performance. Comparisons between traditional machine learning with deep learning approaches, highlight the effectiveness of CNN-LSTM models combined with data augmentation. Despite achieving promising results, most of our experiments indicate a challenge with overfitting, which is likely due to the limited number of samples, even after data augmentation and model regularization. Thus, future work may focus on collecting additional data to address this issue. Abjad-Kids will be publicly available. We hope that Abjad-Kids enrich children representation in speech dataset, and be a good resource for future research in Arabic speech classification for kids.

语音识别儿童语音阿拉伯语教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。