构建普诺克丘亚语语音数据集,推动弱势语言语音技术发展
Quechua Speech Datasets in Common Voice: The Case of Puno Quechua
- 以普诺克丘亚语为案例,通过社区众包收集朗读与自发语音数据
- 共收集191.1小时克丘亚语语音(86%已验证),其中普诺语占12小时(77%验证)
- 关注原住民数据主权,提出兼顾技术与伦理的研究路径
资源匮乏语言如克丘亚语面临语音数据与技术资源短缺,制约其语音技术发展。为应对这一问题,Common Voice提供了开放、社区驱动的语音数据集建设机会。本文研究克丘亚语在Common Voice中的整合现状,详细分析17种克丘亚语,以普诺克丘亚语(ISO 639-3: qxp)为典型案例,涵盖语言上线流程与语料库采集,包括朗读与自发语音数据。结果表明,Common Voice目前已收录191.1小时克丘亚语语音数据(86%已验证),其中普诺克丘亚语贡献12小时(77%已验证),凸显该平台的巨大潜力。本文进一步提出针对技术挑战与社区参与伦理的科研议程,致力于推动包容性语音技术发展,实现弱势语言社区的数字赋能。
原文摘要 · Abstract (English)
Under-resourced languages, such as Quechuas, face data and resource scarcity, hindering their development in speech technology. To address this issue, Common Voice presents a crucial opportunity to foster an open and community-driven speech dataset creation. This paper examines the integration of Quechua languages into Common Voice. We detail the current 17 Quechua languages, presenting Puno Quechua (ISO 639-3: qxp) as a focused case study that includes language onboarding and corpus collection of both reading and spontaneous speech data. Our results demonstrate that Common Voice now hosts 191.1 hours of Quechua speech (86\% validated), with Puno Quechua contributing 12 hours (77\% validated), highlighting the Common Voice's potential. We further propose a research agenda addressing technical challenges, alongside ethical considerations for community engagement and indigenous data sovereignty. Our work contributes towards inclusive voice technology and digital empowerment of under-resourced language communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。