用音频学习隐含音乐风格,让乐谱生成更符合真实演奏风格。
Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping

- 通过跨模态对齐,从音频中提取风格特征并作用于乐谱生成。
- 在钢琴编曲、风格迁移和音视频检索任务上显著提升风格一致性。
- 适合音乐生成、风格迁移研究者及数字音乐创作用户。
音乐风格是什么?尽管常以“摇摆”“古典”或“情感丰富”等文本标签描述,但其真实内涵仍隐含于具体音乐实例中。本文提出一种跨模态框架,从原始音频中学习隐含音乐风格,并应用于符号化音乐生成。受BLIP-2启发,模型利用查询变换器(Q-Former)从大型预训练音频语言模型中提取风格表征,并用于条件化符号语言模型生成钢琴编曲。采用两阶段训练策略:先通过对比学习对齐听觉风格与符号表达,再通过生成建模完成编曲。模型可联合条件于旋律谱(内容)和参考音频样本(风格),实现可控且风格忠实的编曲生成。实验表明,该方法在钢琴伴奏生成、风格迁移和音视频到MIDI检索任务中均取得显著提升,有效增强风格感知对齐与音乐质量。
原文摘要 · Abstract (English)
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation. Inspired by BLIP-2, our model leverages a Querying Transformer (Q-Former) to extract style representations from a large, pre-trained audio language model (LM), and further applies them to condition a symbolic LM for generating piano arrangements. We adopt a two-stage training strategy: contrastive learning to align auditory style with symbolic expression, followed by generative modeling for music arrangement. Our model generates piano performances jointly conditioned on a lead sheet (content) and a reference audio example (style), enabling controllable and stylistically faithful arrangement. Experiments demonstrate the effectiveness of our approach in piano cover generation, style transfer, and audio-to-MIDI retrieval, achieving substantial improvements in style-aware alignment and music quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。