构建首个覆盖22种印度语的超大规模多说话人语音合成数据集
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
- 用去噪模型跨语言增强低质录音,生成高质量语音数据
- 数据集含1704小时语音、10496名说话者,覆盖22种印度语言
- 首次评估印度语音的零样本/少样本泛化能力,适合多语言研究者
近期文本到语音(TTS)合成进展表明,基于大规模网络数据训练的模型可生成高度自然的声音。然而,由于缺乏如LibriVox或YouTube上的高质量人工字幕数据,印度语言相关数据稀缺。为此,我们利用包含自然对话的低质量环境音频识别(ASR)数据集,通过英语训练的降噪与语音增强模型进行跨语言迁移,生成高质量TTS训练数据。由此构建出IndicVoices-R(IV-R),这是首个基于ASR数据集生成的大型多语言印度语TTS数据集,包含1,704小时高质量语音,来自10,496名说话者,覆盖22种印度语言。其语音质量媲美黄金标准数据集(如LJSpeech、LibriTTS、IndicTTS)。我们还提出IV-R基准,首次评估TTS模型在印度语音上的零样本、少样本和多样本说话人泛化能力,确保年龄、性别和风格多样性。实验表明,在结合IndicTTS与我们的IV-R数据集上微调英文预训练模型,相比仅使用IndicTTS,能实现更好的零样本说话人泛化。此外,评估揭示现有模型对印度语音的零样本泛化能力有限,而通过在包含多种语言家族多样说话者的数据上微调,可显著提升。所有数据与代码已开源,发布首个支持全部22种官方印度语言的TTS模型。
原文摘要 · Abstract (English)
Recent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is scarce for Indian languages due to the lack of high-quality, manually subtitled data on platforms like LibriVox or YouTube. To address this gap, we enhance existing large-scale ASR datasets containing natural conversations collected in low-quality environments to generate high-quality TTS training data. Our pipeline leverages the cross-lingual generalization of denoising and speech enhancement models trained on English and applied to Indian languages. This results in IndicVoices-R (IV-R), the largest multilingual Indian TTS dataset derived from an ASR dataset, with 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. IV-R matches the quality of gold-standard TTS datasets like LJSpeech, LibriTTS, and IndicTTS. We also introduce the IV-R Benchmark, the first to assess zero-shot, few-shot, and many-shot speaker generalization capabilities of TTS models on Indian voices, ensuring diversity in age, gender, and style. We demonstrate that fine-tuning an English pre-trained model on a combined dataset of high-quality IndicTTS and our IV-R dataset results in better zero-shot speaker generalization compared to fine-tuning on the IndicTTS dataset alone. Further, our evaluation reveals limited zero-shot generalization for Indian voices in TTS models trained on prior datasets, which we improve by fine-tuning the model on our data containing diverse set of speakers across language families. We open-source all data and code, releasing the first TTS model for all 22 official Indian languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。