为普诺克丘亚语构建首个专用语音识别资源,助力濒危语言数字化保护。
Building Community-Centred NLP Resources for Puno Quechua

- 通过参与式设计收集66小时语音数据,含36小时人工转录验证
- 首次建立普诺克丘亚语语音识别基准,评估多个主流模型性能
- 开源全部数据集与微调模型,支持社区自主使用
濒危语言的数字保护需要由其使用者共同设计的工具与资源。本文首次为普诺克丘亚语(ISO 639-3: qxp)构建专用语音识别资源:(1)收录66小时语音数据,涵盖脚本化与自发性口语,其中36小时经人工转录并验证;(2)建立首个系统性语音识别基准,评估Whisper-base、wav2vec2-base、XLS-R-300M等主流模型,包括持续预训练(CPT)前后表现;(3)公开所有数据集与微调模型,推动社区主导的语言技术发展。
原文摘要 · Abstract (English)
The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR resources for Puno Quechua (ISO 639-3: qxp): (1) the largest speech corpus for any single Quechua variety, consisting in 66 hours of recordings for scripted and spontaneous speech (including 36 hours of manually transcribed and validated data), collected via a participatory design campaign; (2) the first systematic ASR benchmark for Puno Quechua, evaluating state-of-the-art models and fine-tuning Whisper-base, wav2vec2-base, and XLS-R-300M, with and without continued pre-training (CPT); (3) an open release of all datasets and fine-tuned models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。