为巴尔蒂语构建首个语音数据集并训练出低错误率的自动识别系统
BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language

- 基于共用语音数据构建16.8小时巴尔蒂语语音库,使用原生书写系统标注
- 微调Whisper-small使词错误率降至26.74%,字符错误率8.67%
- 开源数据集与模型,适合语言保护与低资源语音研究者使用
我们提出BaltiVoice,一个16.8小时的巴尔蒂语(ISO 639-3: bft)朗读语音语料库,该语言在巴基斯坦吉尔吉特-巴尔蒂斯坦地区使用,此前无公开可用的语音识别资源。语料库包含10,060条经验证的发音,源自Mozilla Common Voice录音,使用原生Nastaliq书写系统。将OpenAI Whisper-small在该数据上微调后,在538条说话人独立的验证集上实现26.74%的词错误率(WER)和8.67%的字符错误率(CER),远低于零样本基线的159.19% WER和152.52% CER。在相同数据上微调Whisper-base得到44.54% WER和15.61% CER,证实模型容量对低资源场景的重要性。数据集、微调模型及实时转录演示已公开发布于HuggingFace。
原文摘要 · Abstract (English)
We present BaltiVoice, a 16.8-hour read-speech corpus for Balti (ISO 639-3: bft), a Tibetic language spoken in Gilgit-Baltistan, Pakistan, with no prior publicly available ASR resources. The corpus contains 10,060 validated utterances in native Nastaliq script, derived from Mozilla Common Voice recordings. Fine-tuning OpenAI Whisper-small yields a Word Error Rate (WER) of 26.74% and a Character Error Rate (CER) of 8.67% on a 538-utterance speaker-disjoint validation set, down from a zero-shot baseline of 159.19% WER and 152.52% CER. A Whisper-base fine-tuned on the same data achieves 44.54% WER and 15.61% CER, confirming that model capacity matters for this low-resource setting. The dataset, fine-tuned model, and a live transcription demo are publicly available on HuggingFace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。