arXiv:2409.13832eess.AScs.CL2024-09NeurIPS被引 65

打造首个全球多技法真实乐谱歌唱数据集,支持全类型歌声任务。

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

  • 收集80.59小时高质歌声,覆盖9种语言20位专业歌手
  • 提供6种演唱技法的逐音素标注与真实乐谱
  • 适配技巧控制、风格迁移等四大任务,免费开源

高质量多任务歌唱数据集稀缺严重制约个性化可控歌唱技术发展。现有数据集普遍存在音质差、语言与歌手多样性不足、缺乏多技法信息与真实乐谱、任务适配性差等问题。为此,我们提出GTSinger——一个大规模、全球性、免费可用、高质量的歌唱语料库,配备真实乐谱,专为所有歌唱任务设计,并附带基准测试。具体包括:(1) 收集80.59小时高保真歌唱音频,构成迄今最大录制歌唱数据集;(2) 20位来自九种主流语言的专业歌手,呈现丰富音色与风格;(3) 提供六种常用演唱技法的可控对比与音素级标注,助力技法建模与控制;(4) 配备真实音乐乐谱,支持实际音乐创作;(5) 歌唱音频附带人工标注的音素-音频对齐、全局风格标签及16.16小时配对语音,满足多种歌唱任务需求。为促进使用,我们开展四项基准实验:技法可控语音合成、技法识别、风格迁移与语音转歌唱。演示地址:http://aaronz345.github.io/GTSingerDemo/。数据与代码见:https://huggingface.co/datasets/AaronZ345/GTSinger 及 https://github.com/AaronZ345/GTSinger。

原文摘要 · Abstract (English)

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and realistic music scores, and poor task suitability. To tackle these problems, we present GTSinger, a large global, multi-technique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks, along with its benchmarks. Particularly, (1) we collect 80.59 hours of high-quality singing voices, forming the largest recorded singing dataset; (2) 20 professional singers across nine widely spoken languages offer diverse timbres and styles; (3) we provide controlled comparison and phoneme-level annotations of six commonly used singing techniques, helping technique modeling and control; (4) GTSinger offers realistic music scores, assisting real-world musical composition; (5) singing voices are accompanied by manual phoneme-to-audio alignments, global style labels, and 16.16 hours of paired speech for various singing tasks. Moreover, to facilitate the use of GTSinger, we conduct four benchmark experiments: technique-controllable singing voice synthesis, technique recognition, style transfer, and speech-to-singing conversion. The demos can be found at http://aaronz345.github.io/GTSingerDemo/. We provide the dataset and the code for processing data and conducting benchmarks at https://huggingface.co/datasets/AaronZ345/GTSinger and https://github.com/AaronZ345/GTSinger.

歌唱合成多语言乐谱生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。