构建61小时多语言语音数据集,提升小语种语音识别性能
EuroSpeech: A Multilingual Speech Corpus
- 从议会录音构建可扩展数据管道,支持长时音频与非逐字转录
- 产出超61000小时语音,19种语言超1000小时,22种超500小时
- 微调现有模型后词错误率降低41.8%,显著改善小语种识别
近期语音处理进展表明,跨语言高质量表现需每种语言具备充足训练数据。尽管现有多语言数据集涵盖众多语言,但多数语言数据量不足,导致模型在多数支持语言上表现不佳。本文提出一种从议会录音构建语音数据集的可扩展流水线,包含鲁棒的媒体检索组件和两阶段对齐算法,能有效处理非逐字转录与长时音频。将该方法应用于22个欧洲议会的录音,共提取超过61,000小时对齐语音片段,实现显著的单语言覆盖:19种语言超过1,000小时,22种语言超过500小时。在现有自动语音识别模型上微调时,词错误率平均降低41.8%,验证了该方法的有效性。
原文摘要 · Abstract (English)
Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often contain insufficient data for most languages. Thus, trained models perform poorly on the majority of the supported languages. Our work addresses this challenge by introducing a scalable pipeline for constructing speech datasets from parliamentary recordings. The proposed pipeline includes robust components for media retrieval and a two-stage alignment algorithm designed to handle non-verbatim transcripts and long-form audio. Applying this pipeline to recordings from 22 European parliaments, we extract over 61k hours of aligned speech segments, achieving substantial per-language coverage with 19 languages exceeding 1k hours and 22 languages exceeding 500 hours of high-quality speech data. We obtain an average 41.8\% reduction in word error rates over baselines when finetuning an existing ASR model on our dataset, demonstrating the usefulness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。