arXiv:2602.14062cs.CLcs.SD2026-02

分析普什图语语音数据集的规模增长与参与不均问题

From Scarcity to Scale: A Release-Level Analysis of the Pashto Common Voice Dataset

  • 按版本追踪普什图语语音数据集演化,揭示其快速扩张趋势
  • 2025年总量达2768.7小时,其中975.89小时已验证可用于训练
  • 贡献者高度集中(基尼系数0.941),性别标注缺失率超40%

大规模公开语音数据集对构建自动语音识别(ASR)系统至关重要,但许多广泛使用语言在公共资源中仍严重不足。普什图语(逾6000万人使用)长期缺乏适合现代ASR开发的大规模开源语音数据。本文对Mozilla Common Voice语料库中的普什图语部分进行版本级分析,聚焦24.0版(2025年12月),并梳理主要版本的趋势。数据量从2023年中旬的1.49小时迅速增长至2025年的2768.7小时,其中975.89小时已验证可用于监督式ASR训练。除规模外,还分析了验证吞吐量、贡献者参与不平等性、人口统计元数据完整性及验证子集中的句子级集中度。结果显示:参与极度集中(基尼系数0.941),年龄分布强烈偏向青年成人,41.97%的音频片段缺少自报性别标签,限制基于元数据的子群体审计。在文本层面,提示重复程度中等:35.88%的唯一句子占验证录音总量的50%,表明结构集中主要由贡献者活动不均驱动,而非少数提示语主导。这些结果为快速增长的低资源语音语料库提供了量化评估,并指明提升数据成熟度的实际优先方向,如增强验证能力与扩大人口多样性参与。

原文摘要 · Abstract (English)

Large, openly licensed speech datasets are essential for building automatic speech recognition (ASR) systems, yet many widely spoken languages remain underrepresented in public resources. Pashto, spoken by more than 60 million people, has historically lacked large-scale openly licensed speech data suitable for modern ASR development. This paper presents a release-level analysis of the Pashto component of the Mozilla Common Voice corpus, focusing on version 24.0 (December 2025) and contextualizing trends across major releases. We document rapid growth from 1.49 recorded hours in mid-2023 to 2,768.7 total hours in 2025, including 975.89 validated hours available for supervised ASR training. Beyond scale, we analyze validation throughput, contributor participation inequality, demographic metadata completeness, and sentence-level concentration in the validated subset. We find that participation is extremely concentrated (Gini = 0.941), age representation is strongly skewed toward young adults, and 41.97\% of clips lack self-reported gender labels, limiting subgroup auditing based on metadata. At the textual level, prompt reuse is moderate: 35.88\% of unique sentences account for 50\% of validated clips, suggesting that structural concentration is driven primarily by uneven contributor activity rather than dominance of a small prompt set. These results provide a quantitative audit of a rapidly scaling low-resource speech corpus and highlight practical priorities for improving dataset maturity, including expanded validation capacity and broader demographic participation.

语音识别低资源语言数据集分析普什图语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。