为非洲低资源语言设计可持续数据治理框架,保障社区权益。
The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages
- 以社区为中心设计数据采集与许可机制,确保本地主导权。
- 构建isiXhosa语音数据集ViXSD,支持语音识别模型训练。
- 适合关注公平数据治理与非洲语言技术的团队使用。
本文提出Esethu框架,一种专为赋能本地社区、实现语言资源公平利益共享而设计的可持续数据治理方案。该框架依托Esethu许可证——一种新型社区导向型数据许可协议。作为概念验证,我们推出了开放源代码的Vuk'uzenzele isiXhosa语音数据集(ViXSD),该数据集由母语isiXhosa使用者朗读的语音构成,并附有人口统计与语言学元数据。实验表明,基于社区驱动的许可与数据治理原则,可有效缓解非洲语言在自动语音识别(ASR)领域的资源缺口,同时保护数据贡献者的权益。本文阐述了指导数据集开发的框架、Esethu许可证的具体条款,介绍了ViXSD的构建方法,并通过ASR实验验证了其在构建和优化isiXhosa语音应用中的可用性。
原文摘要 · Abstract (English)
This paper presents the Esethu Framework, a sustainable data curation framework specifically designed to empower local communities and ensure equitable benefit-sharing from their linguistic resource. This framework is supported by the Esethu license, a novel community-centric data license. As a proof of concept, we introduce the Vuk'uzenzele isiXhosa Speech Dataset (ViXSD), an open-source corpus developed under the Esethu Framework and License. The dataset, containing read speech from native isiXhosa speakers enriched with demographic and linguistic metadata, demonstrates how community-driven licensing and curation principles can bridge resource gaps in automatic speech recognition (ASR) for African languages while safeguarding the interests of data creators. We describe the framework guiding dataset development, outline the Esethu license provisions, present the methodology for ViXSD, and present ASR experiments validating ViXSD's usability in building and refining voice-driven applications for isiXhosa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。