arXiv:2512.11348cs.SD2025-12被引 3

用短语级隐变量扩散模型生成完整多轨音乐,突破传统长序列限制。

PhraseVAE and PhraseLDM: Latent Diffusion for Full-Song Multitrack Symbolic Music Generation

  • 将多声部乐谱压缩为64维短语隐向量,实现高效建模
  • 单次生成128小节(8分钟)完整歌曲,无自回归依赖
  • 仅4500万参数,秒级生成,结构清晰且风格自然

本文提出一种全新的完整歌曲符号音乐生成范式。现有符号模型基于音符属性标记,面临序列过长、上下文受限、长程结构支持弱等问题。我们引入PhraseVAE和PhraseLDM,首个专为完整多轨符号音乐设计的隐变量扩散框架。PhraseVAE将任意长度的多声部音符序列压缩为单一64维短语级隐表示,重建保真度高,构建出结构良好且高效的隐空间。在此基础上,PhraseLDM无需自回归组件,单次生成整首多轨歌曲。系统摒弃小节级顺序建模,支持最多128小节音乐(64 bpm下8分钟),生成作品兼具连贯局部纹理、典型乐器模式与清晰全局结构。模型仅含4500万参数,在秒级内完成全曲生成,同时保持竞争力的音乐质量与生成多样性。结果表明,短语级隐变量扩散为符号音乐生成中的长序列建模提供了有效且可扩展的解决方案。我们希望该工作能推动未来研究超越音符属性标记,将短语级单元作为更有效、更符合音乐意义的建模目标。

原文摘要 · Abstract (English)

This technical report presents a new paradigm for full-song symbolic music generation. Existing symbolic models operate on note-attribute tokens and suffer from extremely long sequences, limited context length, and weak support for long-range structure. We address these issues by introducing PhraseVAE and PhraseLDM, the first latent diffusion framework designed for full-song multitrack symbolic music. PhraseVAE compresses an arbitrary variable-length polyphonic note sequence into a single compact 64-dimensional phrase-level latent representation with high reconstruction fidelity, allowing a well-structured latent space and efficient generative modeling. Built on this latent space, PhraseLDM generates an entire multi-track song in a single pass without any autoregressive components. The system eliminates bar-wise sequential modeling, supports up to 128 bars of music (8 minutes at 64 bpm), and produces complete songs with coherent local texture, idiomatic instrument patterns, and clear global structure. With only 45M parameters, our framework generates a full song within seconds while maintaining competitive musical quality and generation diversity. Together, these results show that phrase-level latent diffusion provides an effective and scalable solution to long-sequence modeling in symbolic music generation. We hope this work encourages future symbolic music research to move beyond note-attribute tokens and to consider phrase-level units as a more effective and musically meaningful modeling target.

音乐生成扩散模型符号音乐隐变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。