arXiv:2512.02201cs.CL2025-12被引 3

打造3000小时多语言语音数据集,助力南非七种语言的语音识别发展

Swivuriso: The South African Next Voices Multilingual Speech Dataset

  • 构建覆盖七种南非语言的3000小时多语言语音数据集
  • 在农业、医疗等场景下实现跨语言语音识别基准测试
  • 注重伦理与数据采集规范,适合非洲语言AI研究者使用

本文介绍了Swivuriso,一个作为非洲下一代声音项目一部分的3000小时多语言语音数据集,旨在支持七种南非语言的自动语音识别(ASR)技术开发与基准测试。该数据集涵盖农业、医疗及通用领域内容,填补了现有ASR数据集的显著空白。文中阐述了数据集设计原则、伦理考量及数据采集流程,并展示了基于该数据集训练/微调的ASR模型基线性能,与相关语言的其他ASR数据集进行了对比。

原文摘要 · Abstract (English)

This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automatic speech recognition (ASR) technologies in seven South African languages. Covering agriculture, healthcare, and general domain topics, Swivuriso addresses significant gaps in existing ASR datasets. We describe the design principles, ethical considerations, and data collection procedures that guided the dataset creation. We present baseline results of training/finetuning ASR models with this data and compare to other ASR datasets for the langauges concerned.

语音识别多语言非洲数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。