arXiv:2603.00961eess.AS2026-03

用歌曲数据提升稀缺语言语音识别性能

Using Songs to Improve Kazakh Automatic Speech Recognition

  • 用3013个歌词语句对构建歌词语音数据集,用于模型微调
  • 混合歌曲与公开语料可使识别错误率降低近一半
  • 适合低资源语言语音系统开发者参考

为低资源语言开发自动语音识别(ASR)系统面临标注语料匮乏的挑战。本概念验证研究探索将歌曲作为非传统但有前景的数据来源用于哈萨克语ASR。我们从36位艺术家的195首歌曲中构建了一个包含3,013个音频-文本对(约4.5小时)的数据集,按歌词行进行切分。以Whisper为基础识别器,我们在七种训练场景下微调模型,涵盖歌曲、Common Voice Corpus(CVC)和FLEURS,并在三个基准上评估:CVC、FLEURS和哈萨克语语音语料库2(KSC2)。结果表明,基于歌曲的微调优于零样本基线。例如,使用歌曲、CVC和FLEURS混合训练的Whisper Large-V3 Turbo在CVC上实现27.6%的归一化词错误率(WER),在FLEURS上为11.8%,而在KSC2上错误率降至39.3%(相较零样本模型的81.2%减半)。尽管性能仍低于在1,100小时的KSC2语料上训练的模型,但表明即使是小规模的歌曲-语音混合也能带来显著的适应性提升。数据集已发布于Hugging Face,供研究使用,采用受控的非商业许可。

原文摘要 · Abstract (English)

Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the scarcity of transcribed corpora. This proof-of-concept study explores songs as an unconventional yet promising data source for Kazakh ASR. We curate a dataset of 3,013 audio-text pairs (about 4.5 hours) from 195 songs by 36 artists, segmented at the lyric-line level. Using Whisper as the base recogniser, we fine-tune models under seven training scenarios involving Songs, Common Voice Corpus (CVC), and FLEURS, and evaluate them on three benchmarks: CVC, FLEURS, and Kazakh Speech Corpus 2 (KSC2). Results show that song-based fine-tuning improves performance over zero-shot baselines. For instance, Whisper Large-V3 Turbo trained on a mixture of Songs, CVC, and FLEURS achieves 27.6% normalised WER on CVC and 11.8% on FLEURS, while halving the error on KSC2 (39.3% vs. 81.2%) relative to the zero-shot model. Although these gains remain below those of models trained on the 1,100-hour KSC2 corpus, they demonstrate that even modest song-speech mixtures can yield meaningful adaptation improvements in low-resource ASR. The dataset is released on Hugging Face for research purposes under a gated, non-commercial licence.

语音识别低资源语言歌曲数据哈萨克语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。