无需标注数据,用拼接+生成模型实现高质量音轨合成
Annotation-Free MIDI-to-Audio Synthesis via Concatenative Synthesis and Generative Refinement
- 先用MIDI拼接生成音频,再用无标注数据训练的扩散模型优化
- 在不依赖标注的情况下,音色和表现力多样性超越现有方法
- 支持通过样本选择精细控制音色,适合音乐制作人使用
基于深度神经网络的MIDI到音频合成方法虽能生成高质量、富有表现力的乐器演奏,但需依赖成对的MIDI-音频标注数据,限制了音色与表现风格的多样性。本文提出CoSaRef,一种无需MIDI-音频配对数据集的合成方法:首先基于输入MIDI通过拼接合成生成音频,再利用在无标注数据上训练的扩散生成模型进行精修。该方法显著提升了音色与表现风格的多样性,并支持通过音频样本选择和额外MIDI设计实现细粒度音色与表现控制,类似传统数字音频工作站功能。实验表明,CoSaRef在无需标注监督的前提下,生成的音轨更具真实感,且在客观与主观评估中均优于基于标注监督的先进音色可控方法。
原文摘要 · Abstract (English)
Recent MIDI-to-audio synthesis methods using deep neural networks have successfully generated high-quality, expressive instrumental tracks. However, these methods require MIDI annotations for supervised training, limiting the diversity of instrument timbres and expression styles in the output. We propose CoSaRef, a MIDI-to-audio synthesis method that does not require MIDI-audio paired datasets. CoSaRef first generates a synthetic audio track using concatenative synthesis based on MIDI input, then refines it with a diffusion-based deep generative model trained on datasets without MIDI annotations. This approach improves the diversity of timbres and expression styles. Additionally, it allows detailed control over timbres and expression through audio sample selection and extra MIDI design, similar to traditional functions in digital audio workstations. Experiments showed that CoSaRef could generate realistic tracks while preserving fine-grained timbre control via one-shot samples. Moreover, despite not being supervised on MIDI annotation, CoSaRef outperformed the state-of-the-art timbre-controllable method based on MIDI supervision in both objective and subjective evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。