用新型架构提升长序列音乐生成效率与质量
Diffusion-based Symbolic Music Generation with Structured State Space Models
- 采用SSM与MFA模块结合,实现线性复杂度下的全局上下文建模
- 在传统中国民歌数据集上超越现有模型,生成质量与效率双优
- 适用于长序列生成任务,对音乐创作与跨领域应用有启发
扩散模型在符号化音乐生成中取得显著进展,但多数方法依赖自注意力机制的Transformer架构,受限于二次计算复杂度,难以扩展至长序列。为此,我们提出基于Mamba的符号音乐扩散模型(SMDIM),融合结构化状态空间模型(SSMs)实现高效全局上下文建模,并引入Mamba-FeedForward-Attention块(MFA),结合了Mamba层的线性复杂度、前馈层的非线性优化以及自注意力的精细控制,平衡可扩展性与音乐表现力。SMDIM达到近似线性复杂度,显著提升长序列处理效率。在包括中国传统民歌数据集FolkDB在内的多个数据集上评估,SMDIM在生成质量与计算效率上均优于当前最优模型。该架构设计亦展现出对多种长序列生成任务的适应潜力,为连贯序列建模提供高效可扩展方案。
原文摘要 · Abstract (English)
Recent advancements in diffusion models have significantly improved symbolic music generation. However, most approaches rely on transformer-based architectures with self-attention mechanisms, which are constrained by quadratic computational complexity, limiting scalability for long sequences. To address this, we propose Symbolic Music Diffusion with Mamba (SMDIM), a novel diffusion-based architecture integrating Structured State Space Models (SSMs) for efficient global context modeling and the Mamba-FeedForward-Attention Block (MFA) for precise local detail preservation. The MFA Block combines the linear complexity of Mamba layers, the non-linear refinement of FeedForward layers, and the fine-grained precision of self-attention mechanisms, achieving a balance between scalability and musical expressiveness. SMDIM achieves near-linear complexity, making it highly efficient for long-sequence tasks. Evaluated on diverse datasets, including FolkDB, a collection of traditional Chinese folk music that represents an underexplored domain in symbolic music generation, SMDIM outperforms state-of-the-art models in both generation quality and computational efficiency. Beyond symbolic music, SMDIM's architectural design demonstrates adaptability to a broad range of long-sequence generation tasks, offering a scalable and efficient solution for coherent sequence modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。