arXiv:2512.19400cs.CL2025-12被引 2

构建160小时巴马拉语语音数据集,提升真实场景下语音识别准确率

Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara

  • 基于马里广播档案构建160小时巴马拉语语音数据集,包含真实口语特征
  • 微调后在两个真实测试集上词错误率降低至37.12%和32.33%
  • 适合关注低资源语言、真实场景语音识别的研究者与开发者

我们提出Kunkado,一个160小时的巴马拉语语音识别数据集,源自马里广播档案,涵盖现代自发性口语,涉及广泛话题。数据集包含代码切换、话语不流畅、背景噪音及多说话人重叠等现实场景特征。我们在33.47小时经人工审核的子集上微调基于Parakeet的模型,并采用实用转录归一化处理以减少数字格式、标签和代码切换标注的变异性。在两个真实测试集上评估,微调后词错误率(WER)从44.47%降至37.12%,另一组从36.07%降至32.33%。人类评估显示,该模型优于使用98小时更清洁、更不真实的语音训练的同架构系统。数据集与模型已公开,以支持主要以口述为主的语言的鲁棒语音识别。

原文摘要 · Abstract (English)

We present Kunkado, a 160-hour Bambara ASR dataset compiled from Malian radio archives to capture present-day spontaneous speech across a wide range of topics. It includes code-switching, disfluencies, background noise, and overlapping speakers that practical ASR systems encounter in real-world use. We finetuned Parakeet-based models on a 33.47-hour human-reviewed subset and apply pragmatic transcript normalization to reduce variability in number formatting, tags, and code-switching annotations. Evaluated on two real-world test sets, finetuning with Kunkado reduces WER from 44.47\% to 37.12\% on one and from 36.07\% to 32.33\% on the other. In human evaluation, the resulting model also outperforms a comparable system with the same architecture trained on 98 hours of cleaner, less realistic speech. We release the data and models to support robust ASR for predominantly oral languages.

语音识别低资源语言真实场景数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。