arXiv:2506.17525cs.CLcs.AI2025-06ACL被引 10

多语言语音数据集存在质量隐患,需重视语言规划与社会语言学意识。

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning

  • 按微观与宏观层面分类数据质量问题,发现弱势语言更易出现宏观缺陷。
  • 以台湾闽南语为例,揭示方言边界与正字法缺失导致数据不可靠。
  • 提出数据构建应融入社区主导的语言复兴策略,适合语言保护研究者。

对三个广泛使用的多语言语音数据集(Mozilla Common Voice 17.0、FLEURS、Vox Populi)的质量审计显示,某些语言的数据存在严重质量问题,可能掩盖下游评估结果的真实表现,制造虚假成功假象。我们将问题分为微观与宏观两类,发现宏观问题在缺乏制度化支持的低资源语言中更为普遍。以台湾闽南语(nan_tw)为例,分析凸显了正字法规范、方言边界界定等主动语言规划的必要性,以及数据构建过程中增强质量控制的重要性。论文最后提出未来数据集开发的指导原则,强调社会语言学意识和语言规划的关键作用,并建议将数据创建过程本身作为社区主导语言复兴与保护的工具。

原文摘要 · Abstract (English)

Our quality audit for three widely used public multilingual speech datasets - Mozilla Common Voice 17.0, FLEURS, and Vox Populi - shows that in some languages, these datasets suffer from significant quality issues, which may obfuscate downstream evaluation results while creating an illusion of success. We divide these quality issues into two categories: micro-level and macro-level. We find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages. We provide a case analysis of Taiwanese Southern Min (nan_tw) that highlights the need for proactive language planning (e.g. orthography prescriptions, dialect boundary definition) and enhanced data quality control in the dataset creation process. We conclude by proposing guidelines and recommendations to mitigate these issues in future dataset development, emphasizing the importance of sociolinguistic awareness and language planning principles. Furthermore, we encourage research into how this creation process itself can be leveraged as a tool for community-led language planning and revitalization.

多语言语音数据质量语言规划社会语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。