arXiv:2606.12282cs.SDcs.LG2026-06

用隐空间对齐让钢琴乐谱生成自然演奏,支持变长输出。

PianoKontext: Expressive Performance Rendering from Deadpan Context

论文配图:PianoKontext: Expressive Performance Rendering from Deadpan Context
图 1 · 摘自论文原文
  • 在音乐隐空间用动态时间规整构建音符与演奏对齐数据
  • 通过DiT块融合编码,实现乐谱到演奏的精准映射
  • 适合想生成自然钢琴演奏的研究者或音乐创作人

表达性演奏生成(EPR)旨在生成受音符序列约束的逼真演奏。然而,现有流匹配音频编辑模型仅处理同步且时长相同的音乐样本,难以理解表达性节奏。本文提出PianoKontext,一种基于预训练Music2Latent模型隐空间的流匹配演奏生成模型,可生成变长钢琴演奏。我们通过合成MIDI乐谱生成无表现力音频,并在隐空间使用动态时间规整(DTW)构建配对训练数据。对齐后的嵌入向量拼接输入DiT模块,有效学习乐谱与演奏间的依赖关系。演示音频可访问:https://realfolkcode.github.io/pianokontext_demo/

原文摘要 · Abstract (English)

Expressive performance rendering (EPR) aims to generate realistic performances constrained on sequences of notes. However, flow matching audio editing models manipulate only synchronized music samples of the same duration, limiting their understanding of expressive timing. We introduce PianoKontext, a flow matching rendering model for classical piano music that generates variable-length performances in the latent space of a pretrained Music2Latent model. We synthesize MIDI scores into deadpan audio and employ Dynamic Time Warping (DTW) in the latent space to construct paired data for training. The aligned embeddings are concatenated in DiT blocks, allowing for a simple and effective learning of the dependencies between the score and performances. Audio samples are available at our demo page: https://realfolkcode.github.io/pianokontext_demo/.

音乐生成流匹配隐空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。