用神经编解码语言模型,让钢琴演奏生成更自然、泛化更强。
MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling
- 将MIDI和音频都转为离散编码,统一建模演奏表现
- 在ATEPP和Maestro数据集上,音质差距降低75%以上
- 适合需要高质量、跨风格钢琴合成的研究与应用
从乐谱生成富有表现力的音频演奏,需同时捕捉乐器声学特征与人类演奏风格。传统方法分两阶段:先生成表现性MIDI,再合成音频。但合成模型常难以泛化到不同MIDI源、音乐风格和录音环境。为此,我们提出MIDI-VALLE,基于原用于零样本个性化语音合成的VALLE框架改进而来。针对演奏MIDI到音频的合成任务,我们优化架构,以参考音频演奏及其对应MIDI作为条件输入。不同于以往依赖钢琴卷帘的TTS系统,MIDI-VALLE将MIDI与音频均编码为离散令牌,实现更一致且鲁棒的钢琴演奏建模。模型通过大规模多样化的钢琴演奏数据集训练,显著提升泛化能力。评估结果显示,相比当前最优基线,MIDI-VALLE在ATEPP和Maestro数据集上弗雷歇音频距离降低超75%;听感测试中获202票,远高于基线的58票,证实其合成质量与跨输入泛化能力的提升。
原文摘要 · Abstract (English)
Generating expressive audio performances from music scores requires models to capture both instrument acoustics and human interpretation. Traditional music performance synthesis pipelines follow a two-stage approach, first generating expressive performance MIDI from a score, then synthesising the MIDI into audio. However, the synthesis models often struggle to generalise across diverse MIDI sources, musical styles, and recording environments. To address these challenges, we propose MIDI-VALLE, a neural codec language model adapted from the VALLE framework, which was originally designed for zero-shot personalised text-to-speech (TTS) synthesis. For performance MIDI-to-audio synthesis, we improve the architecture to condition on a reference audio performance and its corresponding MIDI. Unlike previous TTS-based systems that rely on piano rolls, MIDI-VALLE encodes both MIDI and audio as discrete tokens, facilitating a more consistent and robust modelling of piano performances. Furthermore, the model's generalisation ability is enhanced by training on an extensive and diverse piano performance dataset. Evaluation results show that MIDI-VALLE significantly outperforms a state-of-the-art baseline, achieving over 75% lower Frechet Audio Distance on the ATEPP and Maestro datasets. In the listening test, MIDI-VALLE received 202 votes compared to 58 for the baseline, demonstrating improved synthesis quality and generalisation across diverse performance MIDI inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。