通过分层规划与多轨建模,生成更连贯、结构丰富的完整歌曲。
SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling

- 先生成高阶草图再细化音频,明确整体编排结构。
- 分别建模人声、贝斯、鼓和其他乐器,提升音乐层次感。
- 无需后期优化即可媲美顶尖开源系统,适合音乐创作新手。
现有歌曲生成系统虽能合成逼真音频,但完整歌曲生成仍面临两大挑战:一是缺乏显式的歌曲级编排规划,模型需在生成低级音频细节的同时组织整体结构,常导致段落衔接弱、动态变化不足;二是对不同音乐部分的粗粒度建模掩盖了其独特角色与交互,限制了编排丰富性。本文提出SketchSong,一种分层歌曲生成框架,通过歌曲级草图规划与细粒度多轨建模解决上述问题。时间维度上,SketchSong首先从压缩音频表示中预测紧凑的高阶草图标记序列,再基于这些草图生成音频标记,实现从粗到细的编排规划。轨道维度上,显式建模人声、贝斯、鼓和其他乐器四类轨道,精准捕捉各部分角色与互动。在多个歌曲生成基准上的实验表明,SketchSong在客观指标和主观听感测试中均持续优于基线模型。尽管未采用额外后训练进行偏好优化(如歌词或文本提示对齐),其表现仍可与强大多阶段微调的开源系统比肩,验证了整体设计的有效性。
原文摘要 · Abstract (English)
Recent song generation systems can synthesize realistic audio, yet generating complete songs remains challenging for two reasons. First, explicit song-level arrangement planning remains limited in existing methods, so models often need to organize overall arrangement development while generating low-level audio details. This often leads to incoherence in arrangements, such as weak section transitions and limited dynamic progression. Second, coarse modeling of different musical parts obscures their distinct roles and interactions, limiting arrangement richness of generated songs. In this paper, we present SketchSong, a hierarchical song generation framework that addresses these issues through song-level sketch planning and fine-grained multi-track modeling. Along the temporal dimension, SketchSong first predicts a compact sequence of high-level sketch tokens derived from compressed audio representations, and then generates audio tokens conditioned on these sketches. This coarse-to-fine process gives the model an explicit arrangement plan before detailed audio generation. Along the track dimension, SketchSong explicitly models four tracks, i.e., vocals, bass, drums and other instruments. This enables the model to capture the roles and interactions of different musical parts more precisely. Experiments on song generation benchmarks show that SketchSong consistently outperforms our baseline on both objective metrics and human listening tests. Despite not employing additional post-training for preference optimization such as lyrics and text-prompt alignments, SketchSong achieves competitive results against strong, post-trained open-source systems, demonstrating the effectiveness of our overall design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。