为新语音识别数据集Loquacious提供配套工具与实验验证。
Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
- 构建了n-gram语言模型、音素转换模型及发音词典。
- 在多种语音识别架构上验证数据集有效性。
- 适合研究语音识别泛化与跨领域挑战的学者。
最近发布的Loquacious数据集旨在取代LibriSpeech或TED-Lium等主流英语自动语音识别(ASR)数据集。其主要目标是在多个声学与语言领域中提供明确定义的训练与测试划分,并采用适合学术界和产业界的开放许可。为促进该数据集的基准测试与可用性,本文提供了额外资源:n-gram语言模型、音素到音标(G2P)模型及发音词典,均公开可访问。利用这些资源,我们在多种ASR架构上进行了实验,涵盖不同标签单元与网络结构。初步结果表明,Loquacious数据集为语音识别中的常见挑战提供了有价值的分析案例。
原文摘要 · Abstract (English)
The recently published Loquacious dataset aims to be a replacement for established English automatic speech recognition (ASR) datasets such as LibriSpeech or TED-Lium. The main goal of the Loquacious dataset is to provide properly defined training and test partitions across many acoustic and language domains, with an open license suitable for both academia and industry. To further promote the benchmarking and usability of this new dataset, we present additional resources in the form of n-gram language models (LMs), a grapheme-to-phoneme (G2P) model and pronunciation lexica, with open and public access. Utilizing those additional resources we show experimental results across a wide range of ASR architectures with different label units and topologies. Our initial experimental results indicate that the Loquacious dataset offers a valuable study case for a variety of common challenges in ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。