PerceiverS用多尺度注意力生成长且有表现力的乐谱。
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
- 通过分段与多尺度注意力,同时捕捉音乐结构与演奏细节。
- 在Maestro数据集上生成音乐更连贯多样,结构一致且富有变化。
- 适合想生成长篇复杂乐曲的研究者和作曲家。
基于AI的音乐生成近年来取得显著进展,但生成兼具长时结构与表现力的符号化音乐仍是重大挑战。本文提出PerceiverS(Segmentation and Scale)架构,通过有效分段与多尺度注意力机制,同时学习长期结构依赖与短期表现细节。该方法结合跨注意力与自注意力,在多尺度设置下捕捉长程音乐结构的同时保留演奏细微差别。模型在Maestro数据集上评估,生成音乐展现出更高的连贯性与多样性,具有结构一致性与表现力变化特征。项目演示与生成样本可通过https://perceivers.github.io访问。
原文摘要 · Abstract (English)
AI-based music generation has made significant progress in recent years. However, generating symbolic music that is both long-structured and expressive remains a significant challenge. In this paper, we propose PerceiverS (Segmentation and Scale), a novel architecture designed to address this issue by leveraging both Effective Segmentation and Multi-Scale attention mechanisms. Our approach enhances symbolic music generation by simultaneously learning long-term structural dependencies and short-term expressive details. By combining cross-attention and self-attention in a Multi-Scale setting, PerceiverS captures long-range musical structure while preserving performance nuances. The proposed model has been evaluated using the Maestro dataset and has demonstrated improvements in generating coherent and diverse music, characterized by both structural consistency and expressive variation. The project demos and the generated music samples can be accessed through the link: https://perceivers.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。