研究低资源口语语言语音标注成本,发现每小时语音需30-36小时人工劳动。
Cost Analysis of Human-corrected Transcription for Predominately Oral Languages
- 通过实地与实验室对比,评估母语者对巴马拉语语音的转录耗时。
- 平均每小时语音需30小时(实验室)至36小时(实地)人工标注时间。
- 为低资源口语语言数据构建提供可量化的成本参考,适合资源规划者。
为低资源语言创建语音数据集是关键但尚未充分理解的挑战,尤其在人力成本方面。本文针对马里巴马拉语(一种曼丁语系的低识字率口语语言)开展为期一个月的实地研究,由十名母语者参与,分析自动语音识别生成的53小时巴马拉语语音转录稿的修正过程。研究发现,在实验室条件下,每小时语音需平均30小时人工劳动完成高质量标注;在实地条件下则需36小时。该结果为具有相似特征的大量低资源语言构建NLP资源提供了基准与实用洞见。
原文摘要 · Abstract (English)
Creating speech datasets for low-resource languages is a critical yet poorly understood challenge, particularly regarding the actual cost in human labor. This paper investigates the time and complexity required to produce high-quality annotated speech data for a subset of low-resource languages, low literacy Predominately Oral Languages, focusing on Bambara, a Manding language of Mali. Through a one-month field study involving ten transcribers with native proficiency, we analyze the correction of ASR-generated transcriptions of 53 hours of Bambara voice data. We report that it takes, on average, 30 hours of human labor to accurately transcribe one hour of speech data under laboratory conditions and 36 hours under field conditions. The study provides a baseline and practical insights for a large class of languages with comparable profiles undertaking the creation of NLP resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。