arXiv:2409.06029cs.SDcs.AI2024-09NeurIPS被引 32

根据歌词生成带人声和伴奏的完整歌曲,支持灵活控制音色。

SongCreator: Lyrics-based Universal Song Generation

  • 设计双序列语言模型捕捉人声与伴奏信息
  • 在8项任务中达顶尖或领先表现,歌词转歌曲效果显著提升
  • 可通过音频提示独立控制人声与伴奏音色,适合音乐创作应用

音乐是人类文化的重要组成部分,体现智能与创造力,其中歌曲尤为关键。尽管已有研究探索了歌唱声、人声编排与乐器编排等方向,但根据歌词生成包含人声与伴奏的完整歌曲仍是重大挑战,限制了音乐生成模型的实际应用。为此,我们提出SongCreator,一个专为此难题设计的歌曲生成系统。该模型采用两种创新设计:精心构建的双序列语言模型(DSLM)以捕捉人声与伴奏信息;一系列注意力掩码策略使模型能理解、生成与编辑歌曲,通过特定掩码适配多种歌曲相关生成任务。大量实验表明,SongCreator在全部八项任务中均达到最先进或有竞争力的表现,尤其在歌词转歌曲和歌词转人声任务上显著优于此前工作。此外,模型可通过不同音频提示独立控制生成歌曲中的人声与伴奏音质,展现出良好的实用性。样本已公开于https://thuhcsi.github.io/SongCreator/。

原文摘要 · Abstract (English)

Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating songs with both vocals and accompaniment given lyrics remains a significant challenge, hindering the application of music generation models in the real world. In this light, we propose SongCreator, a song-generation system designed to tackle this challenge. The model features two novel designs: a meticulously designed dual-sequence language model (DSLM) to capture the information of vocals and accompaniment for song generation, and a series of attention mask strategies for DSLM, which allows our model to understand, generate and edit songs, making it suitable for various songrelated generation tasks by utilizing specific attention masks. Extensive experiments demonstrate the effectiveness of SongCreator by achieving state-of-the-art or competitive performances on all eight tasks. Notably, it surpasses previous works by a large margin in lyrics-to-song and lyrics-to-vocals. Additionally, it is able to independently control the acoustic conditions of the vocals and accompaniment in the generated song through different audio prompts, exhibiting its potential applicability. Our samples are available at https://thuhcsi.github.io/SongCreator/.

歌曲生成双序列模型音频控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。