构建首个阿坎语障碍语音数据集,助力低资源语言语音识别发展
Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language
- 收集4类阿坎语障碍语音,涵盖口吃、脑瘫、唇腭裂和中风后失语
- 包含50.01小时音频、转写文本与说话人、环境等元数据
- 为低资源语言障碍语音识别研究提供关键数据支持
缺乏障碍语音数据制约了包容性语音技术的发展,尤其在低资源语言如阿坎语中更为突出。为此,本研究构建了一个由母语阿坎语障碍者提供的语音语料库。数据集包含50.01小时音频,覆盖口吃、脑瘫、唇腭裂及中风后失语四类障碍语音。录音在受控监督环境下进行,参与者用自身语言描述预选图片。数据集包含音频、转写文本及说话人人口统计、障碍类型、录音环境与设备等元数据,旨在支持低资源障碍语音自动识别系统与辅助语音技术研究。
原文摘要 · Abstract (English)
The lack of impaired speech data hinders advancements in the development of inclusive speech technologies, particularly in low-resource languages such as Akan. To address this gap, this study presents a curated corpus of speech samples from native Akan speakers with speech impairment. The dataset comprises of 50.01 hours of audio recordings cutting across four classes of impaired speech namely stammering, cerebral palsy, cleft palate, and stroke induced speech disorder. Recordings were done in controlled supervised environments were participants described pre-selected images in their own words. The resulting dataset is a collection of audio recordings, transcriptions, and associated metadata on speaker demographics, class of impairment, recording environment and device. The dataset is intended to support research in low-resource automatic disordered speech recognition systems and assistive speech technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。