将零样本歌曲生成模型改造为多任务编辑工具,支持歌词、人声、伴奏的灵活修改。
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
- 基于语言模型架构,融合分词器与扩散生成器实现端到端编辑
- 支持段落级和音轨级修改,可生成完整段落或分离人声伴奏
- 首个将编辑能力融入歌曲生成的语言模型,适合音乐创作与制作人群
新型生成建模范式,尤其是音频语言模型,显著推进了歌曲生成领域。尽管当前顶尖模型能够同时生成长达数分钟的人声与伴奏,但对已有歌曲的局部调整或编辑研究仍较薄弱,限制了创作灵活性。本文提出SongEditor,首个将编辑能力引入语言模型歌曲生成的方法,支持段落级与音轨级修改。SongEditor具备调整歌词、人声与伴奏的能力,也可从零生成歌曲。其核心组件包括音乐分词器、自回归语言模型与扩散生成器,可生成完整段落、遮蔽歌词,甚至分离人声与背景音乐。大量实验表明,所提方法在端到端歌曲编辑任务中表现卓越,客观与主观评估均验证其有效性。音频样例见https://cypress-yang.github.io/SongEditor_demo/
原文摘要 · Abstract (English)
The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about partial adjustments or editing of existing songs is still underexplored, which allows for more flexible and effective production. In this paper, we present SongEditor, the first song editing paradigm that introduces the editing capabilities into language-modeling song generation approaches, facilitating both segment-wise and track-wise modifications. SongEditor offers the flexibility to adjust lyrics, vocals, and accompaniments, as well as synthesizing songs from scratch. The core components of SongEditor include a music tokenizer, an autoregressive language model, and a diffusion generator, enabling generating an entire section, masked lyrics, or even separated vocals and background music. Extensive experiments demonstrate that the proposed SongEditor achieves exceptional performance in end-to-end song editing, as evidenced by both objective and subjective metrics. Audio samples are available in https://cypress-yang.github.io/SongEditor_demo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。