用神经音频编码器将鼓谱转为真实鼓声,提升音乐表现力。
Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs

- 通过Transformer预测音频编码器的离散码,实现鼓谱到音频的生成
- 在E-GMD数据集上验证,不同编码器对音质影响显著
- 适合音乐生成、编曲工具开发者参考
从符号化表示直接生成逼真鼓声是音乐感知与机器学习交叉领域的挑战任务。本文提出一种系统,将包含微时序和力度信息的表达性鼓谱(鼓格)转换为鼓音频,方法是预测神经音频编码器的离散码。采用基于Transformer的模型将鼓格输入映射为编码器标记序列,再通过预训练的编码器解码器生成波形音频。实验对比了EnCodec、DAC和X-Codec三种先进神经音频编码器,评估其对鼓声合成质量的影响。系统在大规模人类鼓演奏数据集Expanded Groove MIDI Dataset(E-GMD)上训练与评估,使用客观指标衡量生成音频的保真度与音乐对齐性。结果表明,编码器标记预测是鼓格转音频的有效路径,并为打击乐合成中音频分词器的选择提供了实用洞见。
原文摘要 · Abstract (English)
Generating realistic drum audio directly from symbolic representations is a challenging task at the intersection of music perception and machine learning. We propose a system that transforms an expressive drum grid, a time-aligned MIDI representation with microtiming and velocity information, into drum audio by predicting discrete codes of a neural audio codec. Our approach uses a Transformer-based model to map the drum grid input to a sequence of codec tokens, which are then converted to waveform audio via a pre-trained codec decoder. We experiment with multiple state-of-the-art neural codecs, namely EnCodec, DAC, and X-Codec, to assess how the choice of audio representation impacts the quality of the generated drums. The system is trained and evaluated on the Expanded Groove MIDI Dataset, E-GMD, a large collection of human drum performances with paired MIDI and audio. We evaluate the fidelity and musical alignment of the generated audio using objective metrics. Overall, our results establish codec-token prediction as an effective route for drum grid-to-audio generation and provide practical insights into selecting audio tokenizers for percussive synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。