arXiv:2607.20253cs.SDcs.AI2026-07被引 1

一站式生成完整歌曲,支持歌词作曲、无伴奏创作与风格翻唱。

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

论文配图:Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
图 1 · 摘自论文原文
  • 分层自回归建模+流匹配渲染,实现高质量全曲生成
  • 在多语言评测中表现优异,覆盖有声与无伴奏音乐生成
  • 适合音乐生成、创意作曲与跨风格重编场景

本文提出一个统一的歌曲生成框架,可基于歌词、文本描述和音乐属性生成高质量完整音乐。该框架支持三项任务:从文本生成完整歌曲(歌词转歌曲)、生成无演唱的器乐作品,以及在保留旋律的前提下对现有歌曲进行风格重编(翻唱生成)。系统包含四个核心模块:语义感知编码器将音频编码为8个码本的RVQ离散表示;hybird-LM基于这些离散令牌执行分层自回归建模以完成全曲生成;FullDiT在连续VAE隐空间中进行全曲流匹配,条件于编码令牌、歌词与文本描述以提升音质;对于翻唱生成,旋律模块从参考音频中提取并离散化旋律线索,指导生成同时保留原旋律。此外,我们研究了DPO、GRPO与OPD等基于奖励的后训练策略用于hybird-LM,并采用基于流的GRPO优化FullDiT以增强音乐性与渲染质量。在多语言自动基准测试及Artificial Analysis Music with Vocals排行榜上的实验结果表明,该框架在各项评估设置中均达到具有竞争力的表现。

原文摘要 · Abstract (English)

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content. Architecturally, our system consists of four main components: a semantic-aware tokenizer, hybird-LM, FullDiT, and a two-level melody module. The tokenizer encodes audio into 8-codebook RVQ tokens for efficient discrete music representation. Based on these tokens, hybird-LM performs hierarchical autoregressive audio-token modeling for full-song generation. To improve audio fidelity, FullDiT performs full-song flow matching in a continuous VAE latent space conditioned on codec tokens, lyrics, and text captions. For cover song generation, the melody module extracts and discretizes melody cues from reference audio to guide generation while preserving the original melodic content. Finally, we investigate DPO, GRPO, and OPD as reward-based post-training strategies for hybird-LM and apply flow-based GRPO to FullDiT to improve musicality and rendering quality. Experimental results on a multilingual automatic benchmark, complemented by the Artificial Analysis Music with Vocals leaderboard, show that the proposed framework achieves competitive performance in the evaluated settings.

歌曲生成流匹配分层建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。