构建3000小时跨语言多样歌声数据集,推动歌声合成与转换研究
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
- 从网络音源提取数据,构建覆盖多语言多风格的歌声数据集
- 产出3000小时可用歌声数据,支持多种主流语音模型训练
- 开源先进模型并提供基准测试,适合语音生成与音乐AI研究者
由于缺乏公开可用的大规模、多样化数据集,歌声合成(SVS)和歌声转换(SVC)长期受限。为此,我们提出SingNet,一个大规模、多样且真实场景下的歌声数据集。通过设计数据处理流程,从互联网样本包和歌曲中提取可直接使用的训练数据,构建了涵盖多种语言与风格的3000小时歌声数据。为促进使用并验证其有效性,我们在收集的歌声数据上预训练并开源基于Wav2vec2、BigVGAN和NSF-HiFiGAN的多种前沿模型。同时在自动歌词转录(ALT)、神经声码器和歌声转换(SVC)任务上开展基准实验。音频演示可通过https://singnet-dataset.github.io/查看。
原文摘要 · Abstract (English)
The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present SingNet, an extensive, diverse, and in-the-wild singing voice dataset. Specifically, we propose a data processing pipeline to extract ready-to-use training data from sample packs and songs on the internet, forming 3000 hours of singing voices in various languages and styles. Furthermore, to facilitate the use and demonstrate the effectiveness of SingNet, we pre-train and open-source various state-of-the-art (SOTA) models on Wav2vec2, BigVGAN, and NSF-HiFiGAN based on our collected singing voice data. We also conduct benchmark experiments on Automatic Lyric Transcription (ALT), Neural Vocoder, and Singing Voice Conversion (SVC). Audio demos are available at: https://singnet-dataset.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。