用旋律谱信息生成连贯的长篇伴奏,让作曲者掌控核心旋律。
MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
- 基于人声乐谱的节奏和和弦信息构建稳定音乐路线
- 支持完整歌曲结构,包括无歌词段落的连贯生成
- 适合专业作曲场景,也适用于通用歌词到歌曲生成
端到端的歌词到歌曲模型虽便于普通用户使用,但专业作曲者仍需能保留主旋律控制权的乐谱到歌曲系统。现有乐谱到歌曲方法多限于短片段,难以在长篇生成中保持连贯性,尤其在前奏、桥段等无声段落表现不佳。为此,本文提出MIDI-informed singing accompaniment generation(MIDI-SAG)框架。不同于传统仅依赖音频的模型,MIDI-SAG利用人声乐谱提取的符号化节拍与和弦信息,提供稳定的音乐导航。通过引入结构规划模块,定义时间边界与语义标签,实现有声与无声段落的一致生成。实验表明,该方法可通过专用预训练模块,在单张GPU上实现数据高效训练。结果验证了其在专业乐谱到歌曲及通用歌词到歌曲任务中的潜力。尽管尚属初步探索,但MIDI-SAG为结构化长篇音乐合成提供了可行方向。音频演示已公开,代码将开源至https://composerflow.github.io/web_revealed/。
原文摘要 · Abstract (English)
While end-to-end lyrics-to-song models offer convenience for casual users, professional songwriters require score-to-song systems that allow them to retain authorship over the core melody. However, existing score-to-song methods are limited to short-form snippets and fail to maintain coherence in long-form generation, particularly during vocal-silent sections like intros and bridges. To address this long-form bottleneck, we propose MIDI-informed singing accompaniment generation (MIDI-SAG). Unlike conventional audio-only models, MIDI-SAG utilizes symbolic timing and chord information derived from the vocal MIDI to provide a stable musical roadmap. By incorporating structure planning, which defines temporal boundaries and semantic labels, our framework facilitates consistent generation across both vocal and non-vocal sections. We demonstrate the feasibility of this compositional pipeline by leveraging specialized pre-trained modules, enabling data-efficient training on a single GPU. Our experiments show the potential of this approach for both professional score-to-song and general lyrics-to-song tasks. While an early exploration, MIDI-SAG suggests a promising direction for structured, long-form music synthesis. Audio demos are available, and the code will be open-sourced at https://composerflow.github.io/web_revealed/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。