发布超8972小时葡萄牙语播客数据集,助力语音识别与合成技术发展
Tagarela - A Portuguese speech dataset from podcasts
- 从播客中收集8972小时葡萄牙语音频,专为语音识别和语音合成训练设计
- 采用混合转录策略,结合预训练模型与高保真接口,保证转录准确率
- 数据集可公开获取,适合研究者构建更自然、鲁棒的葡萄牙语语音系统
尽管语音处理技术取得显著进展,葡萄牙语仍因缺乏公开、大规模、高质量的数据集而资源匮乏。为弥补这一空白,我们推出名为TAGARELA的新数据集,包含超过8,972小时的播客音频,专门用于训练自动语音识别(ASR)和文本到语音(TTS)模型。其规模可与英语的GigaSpeech(10kh)相媲美,支持训练最先进的葡萄牙语模型。为确保数据质量,该语料库经过音频预处理,并采用混合策略进行转录:利用先前在高保真转录上训练的ASR模型,结合专有API生成的高质量转录结果,保障初始准确性。最后,我们仅使用该数据集训练了ASR和TTS模型并评估其性能,证明其在推动更鲁棒、更自然的葡萄牙语语音技术方面的潜力。该数据集已公开发布,网址为https://freds0.github.io/TAGARELA/,以促进相关技术的发展。
原文摘要 · Abstract (English)
Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over 8,972 hours of podcast audio, specifically curated for training automatic speech recognition (ASR) and text-to-speech (TTS) models. Notably, its scale rivals English's GigaSpeech (10kh), enabling state-of-the-art Portuguese models. To ensure data quality, the corpus was subjected to an audio pre-processing pipeline and subsequently transcribed using a mixed strategy: we applied ASR models that were previously trained on high-fidelity transcriptions generated by proprietary APIs, ensuring a high level of initial accuracy. Finally, to validate the effectiveness of this new resource, we present ASR and TTS models trained exclusively on our dataset and evaluate their performance, demonstrating its potential to drive the development of more robust and natural speech technologies for Portuguese. The dataset is released publicly, available at https://freds0.github.io/TAGARELA/, to foster the development of robust speech technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。