让鼓谱精准生成指定音色的鼓声音频,创作更自由。
Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis

- 用参考音频风格控制鼓谱生成,实现音色可调。
- 在多个指标上表现优异,节奏对齐与节拍连续性佳。
- 适合音乐制作人快速生成符合需求的鼓循环。
当前数字音乐制作中的鼓循环生成方法(如单次采样或重采样)通常需要创作者投入大量精力。尽管近期生成模型在保真度和文本遵循方面表现良好,但缺乏对鼓乐创作所需的精确控制。现有符号到音频的研究多聚焦于单一音色乐器,未能解决多声部、打击乐器合成的挑战。本文提出「Break-the-Beat!」,一种能够将鼓谱转换为具有参考音频音色的鼓声音频的模型。该模型通过微调预训练文本到音频模型,并引入内容编码器与混合条件机制构建。为支持此目标,我们从现有鼓音频数据集中构建了一个新的配对目标-参考鼓音频数据集。实验表明,该模型能生成高质量鼓音频,准确遵循高分辨率鼓谱,在音频质量、节奏对齐与节拍连续性等指标上表现优秀。这为音乐制作者提供了一种新的可控创作工具。演示页面:https://ik4sumii.github.io/break-the-beat/
原文摘要 · Abstract (English)
Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial efforts of creators. While recent generative models achieve high fidelity and adhere to text, they lack the specific control needed for such a task. Existing symbolic-to-audio research often focuses on single, tonal instruments, leaving the challenge of polyphonic, percussive drum synthesis unaddressed. We address this gap by introducing ``Break-the-Beat!,'' a model capable of rendering a drum MIDI with the timbre of a reference audio. It is built by fine-tuning a pre-trained text-to-audio model with our proposed content encoder and a effective hybrid conditioning mechanism. To enable this, we construct a new dataset of paired target-reference drum audio from existing drum audio datasets. Experiments demonstrate that our model generates high-quality drum audio that follows high-resolution drum MIDI, achieving strong performance across metrics of audio quality, rhythmic alignment, and beat continuity. This offer producers a new, controllable tool for creative production. Demo page: https://ik4sumii.github.io/break-the-beat/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。