用新架构精准生成有民族调式感的古风音乐
MusicMamba: A Dual-Feature Modeling Approach for Generating Chinese Traditional Music with Modal Precision
- 结合Mamba与Transformer,捕捉旋律长程依赖和全局结构
- 在11小时古乐数据集上生成旋律更符合传统调式特征
- 适合做非遗音乐数字化或游戏古风配乐的开发者
近年来,深度学习显著推动了MIDI领域发展,使音乐生成成为人工智能的重要应用。然而,现有研究多聚焦西方音乐,在生成中国民族音乐时难以准确捕捉调式特征与情感表达。为此,本文提出双特征建模模块,融合Mamba块的长程依赖建模能力与Transformer块的全局结构捕捉能力,并引入双向Mamba融合层,通过双向扫描整合局部细节与全局结构,增强复杂序列建模。基于该架构,提出REMI-M表示法,更精准地表征与生成旋律中的调式信息。为支持研究,构建了包含多种风格的高质量中文传统音乐数据集FolkDB,总时长超过11小时。实验表明,所提方法在生成具有中国传统音乐特征的旋律方面表现优异,为音乐生成提供了新且有效的解决方案。
原文摘要 · Abstract (English)
In recent years, deep learning has significantly advanced the MIDI domain, solidifying music generation as a key application of artificial intelligence. However, existing research primarily focuses on Western music and encounters challenges in generating melodies for Chinese traditional music, especially in capturing modal characteristics and emotional expression. To address these issues, we propose a new architecture, the Dual-Feature Modeling Module, which integrates the long-range dependency modeling of the Mamba Block with the global structure capturing capabilities of the Transformer Block. Additionally, we introduce the Bidirectional Mamba Fusion Layer, which integrates local details and global structures through bidirectional scanning, enhancing the modeling of complex sequences. Building on this architecture, we propose the REMI-M representation, which more accurately captures and generates modal information in melodies. To support this research, we developed FolkDB, a high-quality Chinese traditional music dataset encompassing various styles and totaling over 11 hours of music. Experimental results demonstrate that the proposed architecture excels in generating melodies with Chinese traditional music characteristics, offering a new and effective solution for music generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。