用降噪与智能筛选提升众包语音质量,让低资源语言也能训练出好语音模型。
Enhancing Crowdsourced Audio for Text-to-Speech Models
- 对众包语音数据先降噪再用NISQA模型自动筛选高质量样本。
- 处理后语音的UTMOS评分提升0.4,显著优于原始数据。
- 适合想用低成本众包数据训练语音模型的研究者和开发者。
高质量音频数据是训练鲁棒文本到语音(TTS)模型的关键,但常限制于偶然或众包数据集的使用。本文针对以加泰罗尼亚语为主的CommonVoice数据集——一个因噪声和变异性而著称的众包语料库——提出一种降噪流水线,包含音频增强与选择性过滤策略。我们开发了基于非侵入式语音质量评估(NISQA)模型的自动过滤机制,用于识别并保留增强后的高质量样本。为验证该方法的有效性,我们在处理后的数据集上训练了一个先进的基于扩散模型的TTS系统。结果显示,相比未增强的基线数据集,其UTMOS得分提升了0.4。该方法在扩展众包数据在TTS应用中的适用性方面展现出潜力,尤其适用于加泰罗尼亚语等中低资源语言。
原文摘要 · Abstract (English)
High-quality audio data is a critical prerequisite for training robust text-to-speech models, which often limits the use of opportunistic or crowdsourced datasets. This paper presents an approach to overcome this limitation by implementing a denoising pipeline on the Catalan subset of Commonvoice, a crowd-sourced corpus known for its inherent noise and variability. The pipeline incorporates an audio enhancement phase followed by a selective filtering strategy. We developed an automatic filtering mechanism leveraging Non-Intrusive Speech Quality Assessment (NISQA) models to identify and retain the highest quality samples post-enhancement. To evaluate the efficacy of this approach, we trained a state of the art diffusion-based TTS model on the processed dataset. The results show a significant improvement, with an increase of 0.4 in the UTMOS Score compared to the baseline dataset without enhancement. This methodology shows promise for expanding the utility of crowdsourced data in TTS applications, particularly for mid to low resource languages like Catalan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。