用流模型实现情感可控的钢琴演奏生成,效果接近真人。
SyMuPe: Affective and Controllable Symbolic Music Performance
- 基于条件流匹配,支持无条件生成与演奏补全
- 在2968小时对齐数据上训练,性能媲美真人录音
- 通过文本和情绪标签控制演奏情感,适合交互应用
情感是音乐创作与感知的核心。本文提出SyMuPe框架,构建可情感表达且可控的符号化钢琴演奏模型。旗舰模型PianoFlow采用条件流匹配,解决多种多掩码演奏补全任务,支持无条件生成与补全。训练使用2,968小时对齐的乐谱与表现性MIDI数据。通过集成钢琴演奏情绪分类器,并以情绪加权的Flan-T5文本嵌入作为条件输入,实现文本与情绪控制。客观与主观评估显示,PianoFlow优于基于Transformer的基线模型,性能接近人类录制并转录的MIDI样本。不同文本条件下的生成样例分析验证了情绪控制的有效性。该模型可融入交互应用,推动更易用、更吸引人的音乐表演系统发展。
原文摘要 · Abstract (English)
Emotions are fundamental to the creation and perception of music performances. However, achieving human-like expression and emotion through machine learning models for performance rendering remains a challenging task. In this work, we present SyMuPe, a novel framework for developing and training affective and controllable symbolic piano performance models. Our flagship model, PianoFlow, uses conditional flow matching trained to solve diverse multi-mask performance inpainting tasks. By design, it supports both unconditional generation and infilling of music performance features. For training, we use a curated, cleaned dataset of 2,968 hours of aligned musical scores and expressive MIDI performances. For text and emotion control, we integrate a piano performance emotion classifier and tune PianoFlow with the emotion-weighted Flan-T5 text embeddings provided as conditional inputs. Objective and subjective evaluations against transformer-based baselines and existing models show that PianoFlow not only outperforms other approaches, but also achieves performance quality comparable to that of human-recorded and transcribed MIDI samples. For emotion control, we present and analyze samples generated under different text conditioning scenarios. The developed model can be integrated into interactive applications, contributing to the creation of more accessible and engaging music performance systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。