arXiv:2409.08103cs.CLcs.SD2024-09

为极度匮乏资源的法语普罗旺斯方言构建语音识别基准,推动低资源语音技术发展。

The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language

  • 基于田野录音构建无标准拼写的法语普罗旺斯方言语音数据集
  • 在5小时标注数据上达到30.4%的最优音素错误率
  • 适合研究低资源语音识别与自监督学习的学者参考

我们提出了一个名为Faetar的自动语音识别基准,旨在推动当前低资源语音识别方法的极限。Faetar是意大利主要使用的法兰克-普罗旺斯语方言,缺乏标准拼写,几乎没有任何现存的文本或语音资源,仅包含该基准中的内容。该语料库来自田野录音,多数为嘈杂环境录制,仅有5小时语音配有对应转录,强制对齐质量参差不齐。此外,语料库还包含额外20小时未标注语音。我们报告了使用最先进的多语言语音基础模型的基线结果,在持续预训练于未标注集的流程下,达到最优30.4%的音素错误率。

原文摘要 · Abstract (English)

We introduce the Faetar Automatic Speech Recognition Benchmark, a benchmark corpus designed to push the limits of current approaches to low-resource speech recognition. Faetar, a Franco-Provençal variety spoken primarily in Italy, has no standard orthography, has virtually no existing textual or speech resources other than what is included in the benchmark, and is quite different from other forms of Franco-Provençal. The corpus comes from field recordings, most of which are noisy, for which only 5 hrs have matching transcriptions, and for which forced alignment is of variable quality. The corpus contains an additional 20 hrs of unlabelled speech. We report baseline results from state-of-the-art multilingual speech foundation models with a best phone error rate of 30.4%, using a pipeline that continues pre-training on the foundation model using the unlabelled set.

语音识别低资源方言建模自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。