构建首个公开的库尔德语北部方言语音基准数据集,支持语音识别与翻译研究。
FLEURS-Kobani: Extending the FLEURS Dataset for Northern Kurdish
- 扩展FLEURS数据集,新增5162条库尔德语北部方言语音
- 语音识别最佳表现:词错误率28.11,字符错误率9.84
- 适合低资源语言语音技术研究者使用
FLEURS提供了100多种语言的多语种语音数据,但未包含北部库尔德语,限制了该语言在自动语音识别(ASR)和语音翻译(S2TT)任务中的基准评估。本文提出FLEURS-Kobani,是首个针对北部库尔德语(ISO 639-3 KMR)的语音扩展数据集。该数据集包含5,162条经验证的语音片段,总计18小时24分钟,由31名母语者录制。它首次为该低资源库尔德语变体提供公共评估基准。作为基线,我们微调Whisper v3-large进行ASR和端到端语音翻译(E2E S2TT)。采用两阶段微调策略(Common Voice→FLEURS-Kobani)获得最佳性能:测试集上词错误率(WER)为28.11,字符错误率(CER)为9.84。对于端到端的库尔德语到英语翻译(KMR→EN),Whisper取得8.68的BLEU分数;同时报告了基于中间目标的翻译结果及级联式翻译设置。FLEURS-Kobani已公开发布,供研究使用,许可协议为CC BY 4.0。
原文摘要 · Abstract (English)
FLEURS offers n-way parallel speech for 100+ languages, but Northern Kurdish is not one of them, which limits benchmarking for automatic speech recognition and speech translation tasks in this language. We present FLEURS-Kobani, a Northern Kurdish (ISO 639-3 KMR) spoken extension of the FLEURS benchmark. The FLEURS-Kobani dataset consists of 5,162 validated utterances, totaling 18 hours and 24 minutes. The data were recorded by 31 native speakers. It extends benchmark coverage to an under-resourced Kurdish variety. As baselines, we fine-tuned Whisper v3-large for ASR and E2E S2TT. A two-stage fine-tuning strategy (Common Voice to FLEURS-Kobani) yields the best ASR performance (WER 28.11, CER 9.84 on test). For E2E S2TT (KMR to EN), Whisper achieves 8.68 BLEU on test; we additionally report pivot-derived targets and a cascaded S2TT setup. FLEURS-Kobani provides the first public Northern Kurdish benchmark for evaluation of ASR, S2TT and S2ST tasks. The dataset is publicly released for research use under a CC BY 4.0 license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。