为濒危语言的语音障碍者构建开源数据集,助力无障碍语音识别
A Cookbook for Community-driven Data Collection of Impaired Speech in LowResource Languages
- 设计社区主导的数据收集手册,指导多方协作采集语音样本
- 建成首个阿坎语障碍语音开源数据集,覆盖多样人群
- 提供工具与方法,适合关注包容性AI的研究者使用
本研究提出一种面向语音障碍者、特别是低资源语言的语音数据收集方法,旨在推动语音识别技术的普惠化。通过制定“操作手册”和培训指南,实现社区驱动的数据采集与模型构建。作为概念验证,研究团队建立了首个阿坎语障碍语音的开源数据集,涵盖来自不同背景的语音障碍参与者。该数据集连同操作手册及开源工具已公开,支持研究人员开发针对特定群体需求的包容性语音识别系统。此外,研究还展示了对开源语音识别模型进行微调后,在阿坎语障碍语音识别上的初步成效。
原文摘要 · Abstract (English)
This study presents an approach for collecting speech samples to build Automatic Speech Recognition (ASR) models for impaired speech, particularly, low-resource languages. It aims to democratize ASR technology and data collection by developing a "cookbook" of best practices and training for community-driven data collection and ASR model building. As a proof-of-concept, this study curated the first open-source dataset of impaired speech in Akan: a widely spoken indigenous language in Ghana. The study involved participants from diverse backgrounds with speech impairments. The resulting dataset, along with the cookbook and open-source tools, are publicly available to enable researchers and practitioners to create inclusive ASR technologies tailored to the unique needs of speech impaired individuals. In addition, this study presents the initial results of fine-tuning open-source ASR models to better recognize impaired speech in Akan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。