arXiv:2606.30642cs.SDcs.AI2026-06被引 3

LeVo 2 用分层建模生成长段歌曲,兼顾音乐连贯性与音轨细节。

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

论文配图:LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training
图 1 · 摘自论文原文
  • 分层建模:先语义规划,再并行生成人声与伴奏
  • 训练策略分离音乐性、可控性与音质优化,提升生成质量
  • 适合需要高质量可控歌曲生成的研究者与创作者

全长歌曲生成需保持连贯性与音乐性,还原人声与伴奏的精细声学特征,并遵循歌词与提示。现有基于语言模型的系统存在结构性权衡:混合标记建模虽保留人声-伴奏协调性,但模糊了轨道特异性细节;双轨道预测改善声学表现,却需更长序列且削弱全局规划。本文提出 LeVo 2,一种融合大语言模型与扩散模型的混合框架,用于可控全长歌曲生成。该框架将权衡问题建模为分层结构:首先由 LeLM 预测混合标记进行语义规划,随后并行预测人声与伴奏标记以实现轨道级细化,最后通过基于扩散的 Music Codec 重建完整波形。本版本核心贡献为美学引导的训练调度:预训练阶段,自动化音乐美学评估框架为大规模数据分配音乐性等级,提供音乐性先验;后续渐进式后训练分别应用 SFT、大规模离线 DPO 及闭环半在线 DPO,独立提升生成质量、可控性与音乐性。模块化扩展则对轨道特异性语言模型进行训练,同时保留已对齐的语义规划器。该策略分离音乐性学习、可控性对齐与声学精炼,缓解优化冲突与静态偏好对的局限。专家听觉测试与客观评估显示,LeVo 2 在六项主观维度上超越开源基线,部分指标接近领先商用系统。消融实验验证了训练策略、美学引导、规模效应及分层架构的有效性。

原文摘要 · Abstract (English)

Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts. Existing language model-based systems face a structural trade-off: mixed-token modeling preserves vocal-instrument coordination but obscures track-specific details, whereas dual-track prediction improves acoustics but requires longer sequences and weakens global planning. We present LeVo 2, a hybrid LLM-Diffusion framework for controllable full-length song generation. LeVo 2 formulates this trade-off as hierarchical modeling: LeLM first predicts mixed tokens for semantic planning, then predicts vocal and accompaniment tokens in parallel for track-specific refinement, while a diffusion-based Music Codec reconstructs full-length waveforms. A central contribution of this extended version is an aesthetics-guided training schedule for alignment. During pre-training, an automated music aesthetic evaluation framework assigns musicality-tier conditions to large-scale data, providing musicality priors before preference alignment. Progressive post-training applies SFT, large-scale offline DPO, and closed-loop semi-online DPO to separately improve generation quality, controllability, and musicality. Modular extension then trains the Track-Specific LM for acoustic refinement while preserving the aligned semantic planner. This schedule separates musicality learning, controllability alignment, and acoustic refinement, mitigating optimization conflict and the limitations of static offline preference pairs. Expert listening tests and objective evaluations show that LeVo 2 outperforms open-source baselines across six subjective dimensions, and approaches leading commercial systems on several listening metrics. Ablations validate the effects of the training strategy, aesthetics guidance, scaling, and hierarchical architecture.

歌曲生成分层建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。