arXiv:2604.10736cs.CLcs.SD2026-04

首个专为爱尔兰语设计的语音识别评测基准,解决文本归一化难题。

BlasBench: An Open Benchmark for Irish Speech Recognition

  • 构建爱尔兰语专用归一化工具,保留变音符号与辅音变化特征
  • 12个模型在两数据集上测试,微软Azure达22.2% WER,开源最佳模型30.65%
  • 跨数据集性能下降33-43点,暴露单数据集评估的局限性

现有多语言基准虽包含爱尔兰语,但未采用爱尔兰语感知的文本归一化,导致可靠的语音识别比较无法实现。本文提出BlasBench,一个开放的评测框架,包含独立的爱尔兰语归一化器,能保留变音符号(fadas)、软化(lenition)和鼻音化(eclipsis);提供可复现的评分机制及每句预测结果。我们以Common Voice ga-IE和FLEURS ga-IE为基准,对12个系统在四种架构家族中进行评测。所有Whisper变体因插入错误导致WER超过100%;Microsoft Azure在Common Voice上达22.2% WER,FLEURS上为57.5%;最佳开源模型Omnilingual ASR 7B分别达到30.65%和39.09%。在Common Voice上微调的模型迁移到FLEURS时性能下降33-43点,而大规模多语言模型仅下降7-10点,揭示了单数据集评估所忽视的泛化差距。

原文摘要 · Abstract (English)

Existing multilingual benchmarks include Irish among dozens of languages but apply no Irish-aware text normalisation, leaving reliable and reproducible ASR comparison impossible. We introduce BlasBench, an open evaluation harness that provides a standalone Irish-aware normaliser preserving fadas, lenition, and eclipsis; a reproducible scoring harness and per-utterance predictions released for all evaluated runs. We pilot this by benchmarking 12 systems across four architecture families on Common Voice ga-IE and FLEURS ga-IE. All Whisper variants exceed 100% WER through insertion-driven hallucination. Microsoft Azure reaches 22.2% WER on Common Voice and 57.5% on FLEURS; the best open model, Omnilingual ASR 7B, reaches 30.65% and 39.09% respectively. Models fine-tuned on Common Voice degrade 33-43 points moving to FLEURS, while massively multilingual models degrade only 7-10 - a generalisation gap that single-dataset evaluation misses.

语音识别多语言爱尔兰语评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。