用三阶段方法生成结构一致的钢琴伴奏,提升流畅度与风格可控性。
Etude: Piano Cover Generation with a Three-Stage Approach -- Extract, strucTUralize, and DEcode
- 分三阶段提取节奏、结构化建模、解码生成,增强结构一致性
- 主观评测显示质量接近人类作曲水平,显著优于现有模型
- 支持风格注入,适合音乐生成与创意编曲场景
钢琴伴奏生成旨在将流行歌曲自动转换为钢琴演奏版本。尽管已有诸多深度学习方法提出,但现有模型常难以保持与原曲的结构一致性,可能源于缺乏节拍感知机制或对复杂节奏模式建模困难。节奏信息至关重要,它定义了结构相似性(如速度、BPM),并直接影响生成音乐的整体质量。本文提出Etude,一种包含提取(Extract)、结构化(strucTUralize)和解码(DEcode)三个阶段的三阶段架构。通过预先提取节奏信息,并采用新颖简化的REMI基础标记化方法,该模型生成的伴奏能有效保持歌曲结构,提升流畅性与音乐动态表现力,并可通过风格注入实现高度可控的生成。主观评估显示,Etude在人类听者中显著优于先前模型,其生成质量已达到接近人类作曲者的水平。
原文摘要 · Abstract (English)
Piano cover generation aims to automatically transform a pop song into a piano arrangement. While numerous deep learning approaches have been proposed, existing models often fail to maintain structural consistency with the original song, likely due to the absence of beat-aware mechanisms or the difficulty of modeling complex rhythmic patterns. Rhythmic information is crucial, as it defines structural similarity (e.g., tempo, BPM) and directly impacts the overall quality of the generated music. In this paper, we introduce Etude, a three-stage architecture consisting of Extract, strucTUralize, and DEcode stages. By pre-extracting rhythmic information and applying a novel, simplified REMI-based tokenization, our model produces covers that preserve proper song structure, enhance fluency and musical dynamics, and support highly controllable generation through style injection. Subjective evaluations with human listeners show that Etude substantially outperforms prior models, achieving a quality level comparable to that of human composers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。