为西非低资源语言巴马拉语构建612小时语音数据集并训练小型模型
Dealing with the Hard Facts of Low-Resource African NLP
- 采集612小时巴马拉语自发语音,半自动标注转录
- 基于数据集训练超紧凑与小型模型,验证性能
- 提供数据/模型/代码开源,强调人工评估重要性
针对低资源语言的语音数据集、模型和评估框架构建仍具挑战,因缺乏可借鉴的实践经验。本文报告了在西非低资源语言巴马拉语中开展的实地语音采集,共获得612小时自发口语数据;采用半自动化方式对数据进行转录标注;基于该数据集训练了多个单语超紧凑及小型模型,并进行了自动与人工评估。研究提出了数据采集、标注及模型设计的实用建议,证明了人工评估的重要性。除主数据集外,还公开了多个评估数据集、模型及代码。
原文摘要 · Abstract (English)
Creating speech datasets, models, and evaluation frameworks for low-resource languages remains challenging given the lack of a broad base of pertinent experience to draw from. This paper reports on the field collection of 612 hours of spontaneous speech in Bambara, a low-resource West African language; the semi-automated annotation of that dataset with transcriptions; the creation of several monolingual ultra-compact and small models using the dataset; and the automatic and human evaluation of their output. We offer practical suggestions for data collection protocols, annotation, and model design, as well as evidence for the importance of performing human evaluation. In addition to the main dataset, multiple evaluation datasets, models, and code are made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。