arXiv:2409.15350eess.AScs.CL2024-09被引 5

首个面向圣保罗口音葡萄牙语的自发语音数据集,可用于提升语音识别准确率。

A Large Dataset of Spontaneous Speech with the Accent Spoken in São Paulo for Automatic Speech Recognition Evaluation

  • 构建401人、239小时的自发口语数据集,专注圣保罗口音
  • 最佳模型在该数据集上词错误率低至24.22%
  • 开源数据与训练代码,支持复现与后续研究

我们发布了一个免费可用的巴西葡萄牙语自发语音语料库,并报告了基于Wav2Vec2-XLSR-53和Distil-Whisper模型在该语料库上微调与训练的初步自动语音识别(ASR)结果。NURC-SP音频语料库包含401名不同说话者(204名女性,197名男性),总计239.30小时已转录的音频记录。据我们所知,这是首个专为葡萄牙语语音识别任务设计的大型圣保罗口音自发语音语料库。本文首先介绍NURC-SP语料库的设计与开发流程,随后详细描述四项ASR实验。实验表明该语料库在语音识别应用中具有良好潜力。具体而言,我们微调了两个版本的Wav2Vec2-XLSR-53模型,使用由Whisper Large-V3模型生成标签的数据训练了一个Distil-Whisper模型,并进一步在本语料库上微调该模型。最佳结果为在NURC-SP语料库上微调的Distil-Whisper模型,词错误率为24.22%,优于微调版Wav2Vec2-XLSR-53模型的33.73%,差距接近10个百分点。为便于实验复现,我们已将NURC-SP音频语料库、预训练模型及训练脚本公开于Hugging Face和Github仓库。

原文摘要 · Abstract (English)

We present a freely available spontaneous speech corpus for the Brazilian Portuguese language and report preliminary automatic speech recognition (ASR) results, using both the Wav2Vec2-XLSR-53 and Distil-Whisper models fine-tuned and trained on our corpus. The NURC-SP Audio Corpus comprises 401 different speakers (204 females, 197 males) with a total of 239.30 hours of transcribed audio recordings. To the best of our knowledge, this is the first large Paulistano accented spontaneous speech corpus dedicated to the ASR task in Portuguese. We first present the design and development procedures of the NURC-SP Audio Corpus, and then describe four ASR experiments in detail. The experiments demonstrated promising results for the applicability of the corpus for ASR. Specifically, we fine-tuned two versions of Wav2Vec2-XLSR-53 model, trained a Distil-Whisper model using our dataset with labels determined by Whisper Large-V3 model, and fine-tuned this Distil-Whisper model with our corpus. Our best results were the Distil-Whisper fine-tuned over NURC-SP Audio Corpus with a WER of 24.22% followed by a fine-tuned versions of Wav2Vec2-XLSR-53 model with a WER of 33.73%, that is almost 10% point worse than Distil-Whisper's. To enable experiment reproducibility, we share the NURC-SP Audio Corpus dataset, pre-trained models, and training recipes in Hugging-Face and Github repositories.

语音识别语音数据集巴西葡语圣保罗口音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。