arXiv:2509.21739cs.SDcs.LG2025-09中稿 · ICASSP 2026被引 2

用扩散模型生成鼓点,速度与精度可调,效果超新高。

Noise-to-Notes: Diffusion-based Generation and Refinement for Automatic Drum Transcription

  • 将鼓点转录视为生成任务,用扩散模型从噪声中生成带力度的鼓事件
  • 引入自适应伪赫伯损失,解决二值起始点与连续力度联合优化难题
  • 融合音乐基础模型特征,显著提升对陌生鼓声的鲁棒性,适合音频生成研究者

传统自动鼓点转录(ADT)被建模为从音频频谱图中预测鼓事件的判别任务。本文将ADT重新定义为条件生成任务,提出噪声到音符(Noise-to-Notes, N2N)框架,利用扩散模型将音频条件下的高斯噪声转化为带有力度信息的鼓事件。该生成式方法具备灵活的速度-精度权衡和强大的补全能力。然而,生成二值起始点与连续力度值对扩散模型构成挑战,为此我们设计了自适应伪赫伯损失以实现有效联合优化。此外,为增强低层频谱特征,我们引入音乐基础模型(MFMs)提取的高层语义特征,提升对域外鼓音频的鲁棒性。实验表明,加入MFMs特征显著提高鲁棒性,且N2N在多个ADT基准上达到新的最先进性能。

原文摘要 · Abstract (English)

Automatic drum transcription (ADT) is traditionally formulated as a discriminative task to predict drum events from audio spectrograms. In this work, we redefine ADT as a conditional generative task and introduce Noise-to-Notes (N2N), a framework leveraging diffusion modeling to transform audio-conditioned Gaussian noise into drum events with associated velocities. This generative diffusion approach offers distinct advantages, including a flexible speed-accuracy trade-off and strong inpainting capabilities. However, the generation of binary onset and continuous velocity values presents a challenge for diffusion models, and to overcome this, we introduce an Annealed Pseudo-Huber loss to facilitate effective joint optimization. Finally, to augment low-level spectrogram features, we propose incorporating features extracted from music foundation models (MFMs), which capture high-level semantic information and enhance robustness to out-of-domain drum audio. Experimental results demonstrate that including MFM features significantly improves robustness and N2N establishes a new state-of-the-art performance across multiple ADT benchmarks.

鼓点转录扩散模型音乐生成特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。