将乐谱直接转为有表现力的钢琴音频,一步到位。
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
- 用Transformer模型+调优音符合成器,端到端生成表现力强的钢琴演奏。
- 在ATEPP数据集上还原真人演奏的细微动态与环境音效。
- 适合音乐生成、智能作曲和音频制作领域研究者使用。
本文提出一个集成系统,可将符号化乐谱直接转化为富有表现力的钢琴演奏音频。通过结合基于Transformer的表现力演奏渲染(EPR)模型与微调的神经MIDI合成器,该方法能从乐谱输入直接生成具有表现力的音频输出。据我们所知,这是首个将无表达控制的乐谱MIDI文件无缝转换为丰富表现力钢琴演奏的系统。我们在ATEPP数据集的子集上进行了实验,采用客观指标与主观听觉测试评估系统性能。结果表明,该系统不仅能准确再现类人表现力,还能捕捉音乐会厅与录音室等不同声学环境的氛围感。此外,生成的音频在保持高质量的同时展现出良好的音乐表现力。
原文摘要 · Abstract (English)
This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rendering (EPR) model with a fine-tuned neural MIDI synthesiser, our approach directly generates expressive audio performances from score inputs. To the best of our knowledge, this is the first system to offer a streamlined method for converting score MIDI files lacking expression control into rich, expressive piano performances. We conducted experiments using subsets of the ATEPP dataset, evaluating the system with both objective metrics and subjective listening tests. Our system not only accurately reconstructs human-like expressiveness, but also captures the acoustic ambience of environments such as concert halls and recording studios. Additionally, the proposed system demonstrates its ability to achieve musical expressiveness while ensuring good audio quality in its outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。