arXiv:2509.19271cs.CLcs.AI2025-09NeurIPS

构建首个沃洛夫语银行语音意图分类数据集,助力低资源语言智能服务发展。

WolBanking77: Wolof Banking Speech Intent Classification Dataset

  • 构建包含9791条文本和4小时语音的沃洛夫语银行领域数据集
  • 在文本与语音模型上取得良好效果,基线F1分数表现优异
  • 适合低资源语言研究、非洲本地化语音助手开发者使用

近年来,意图分类模型取得了显著进展,但以往研究主要集中在高资源语言数据集上,导致低资源语言及文盲率高的地区面临技术鸿沟。以塞内加尔为例,约90%人口使用沃洛夫语,而全国文盲率仍高达42%,沃洛夫语在西非地区有超过1000万人使用。为弥补这一不足,本文提出沃洛夫语银行语音意图分类数据集(WolBanking77),用于学术研究。该数据集目前包含9,791条银行领域文本句子和超过4小时的语音数据。本文在多种基线模型上进行实验,涵盖文本与语音前沿模型。结果表明,该数据集表现良好。同时,文章深入分析了数据内容,报告了在WolBanking77上训练的NLP与自动语音识别(ASR)模型的基线F1分数和词错误率,并进行了模型对比。数据集与代码已开源:https://github.com/abdoukarim/wolbanking77。

原文摘要 · Abstract (English)

Intent classification models have made a significant progress in recent years. However, previous studies primarily focus on high-resource language datasets, which results in a gap for low-resource languages and for regions with high rates of illiteracy, where languages are more spoken than read or written. This is the case in Senegal, for example, where Wolof is spoken by around 90\% of the population, while the national illiteracy rate remains at of 42\%. Wolof is actually spoken by more than 10 million people in West African region. To address these limitations, we introduce the Wolof Banking Speech Intent Classification Dataset (WolBanking77), for academic research in intent classification. WolBanking77 currently contains 9,791 text sentences in the banking domain and more than 4 hours of spoken sentences. Experiments on various baselines are conducted in this work, including text and voice state-of-the-art models. The results are very promising on this current dataset. In addition, this paper presents an in-depth examination of the dataset's contents. We report baseline F1-scores and word error rates metrics respectively on NLP and ASR models trained on WolBanking77 dataset and also comparisons between models. Dataset and code available at: https://github.com/abdoukarim/wolbanking77.

语音识别低资源语言意图分类数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。