发布2.5万小时可商用英语语音数据集,支持真实场景下语音识别研究。
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
- 构建25,000小时可商用英语语音数据集,覆盖多样口音与语料类型。
- 包含超百万说话人,涵盖朗读、即兴、演讲、清晰与嘈杂环境等多种语音。
- 专为学术与工业界联合研发设计,解决现有数据集许可与质量缺陷问题。
自动语音识别(ASR)研究依赖于工业界与学术界共享的通用数据集,以促进对比与评估。尽管LibriSpeech长期作为基准,但其规模有限且仅聚焦于清晰的朗读语音,导致词错率接近零。近期数据集如MOSEL、YODAS、Gigaspeech、OWSM、Libriheavy或People's Speech存在严重局限:许可限制使产业界无法使用、转录不可靠、音频错误或缺乏评估集。本文提出Loquacious Set,一个25,000小时经筛选的可商用英语语音数据集。涵盖数十万说话人,包含丰富口音与广泛语音类型(朗读、即兴、演讲、清晰、嘈杂),专为学术与产业界研究人员在真实场景下构建语音识别系统而设计。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) research is driven by the availability of common datasets between industrial researchers and academics, encouraging comparisons and evaluations. LibriSpeech, despite its long success as an ASR benchmark, is now limited by its size and focus on clean, read speech, leading to near-zero word error rates. More recent datasets, including MOSEL, YODAS, Gigaspeech, OWSM, Libriheavy or People's Speech suffer from major limitations including licenses that researchers in the industry cannot use, unreliable transcriptions, incorrect audio data, or the lack of evaluation sets. This work presents the Loquacious Set, a 25,000-hour curated collection of commercially usable English speech. Featuring hundreds of thousands of speakers with diverse accents and a wide range of speech types (read, spontaneous, talks, clean, noisy), the Loquacious Set is designed to work for academics and researchers in the industry to build ASR systems in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。