用扩散模型分离鼓声并转录,既准又保留可编辑音频。
Separate-and-Detect: Unified Drum Transcription and Stem Generation via Latent Diffusion

- 先分离五类鼓音轨,再检测击打时刻,统一处理流程。
- 在MDB和ENST数据集上,鼓点识别准确率优于现有方法。
- 输出可编辑的鼓音轨,适合音乐制作与后期修改场景。
自动鼓声转录(ADT)通常将混音直接映射为符号化鼓事件,虽有效但丢失了可用于编辑、重混和制作的音频分轨。本文重新审视一种「分离-检测」范式:前端通过五音轨潜空间扩散模型生成踢鼓、军鼓、嗵嗵鼓、踩镲和吊镲的可编辑音轨,后端固定节拍检测器将其转为符号事件。训练时引入仅用于学习的击打检测分支(OB)和音色分支(TB),推理时丢弃。在合成鼓多轨数据上训练,在MDB Drums和ENST-Drums上评估,该方案在整体转录F1上持续优于强基线U-Net分离模型,且在踢鼓与军鼓F1上超过典型端到端ADT系统。消融实验表明,OB带来最稳定的转录提升,而TB调节重建质量、音轨质量和击打检测之间的权衡。结果表明,生成式鼓声解混不仅能作为分离模型,还可作为可解释鼓转录的实用前处理模块。
原文摘要 · Abstract (English)
Automatic Drum Transcription (ADT) is commonly formulated as a direct mapping from a music mixture to symbolic drum events. While effective for transcription, this formulation discards the acoustic stems that are useful for editing, remixing, and production. We revisit an alternative separate-and-detect formulation, where a drum source separation front end first produces five editable drum stems, and a fixed onset detector then converts each stem into symbolic events. The separator is built on a five-stem latent diffusion model that jointly generates kick, snare, toms, hi-hats, and cymbals in a compact VAE latent space. We further study two training-only auxiliary branches--an onset branch (OB) and a timbre branch (TB)--which shape the separator during learning but are discarded at inference. Trained on synthetic drum multitracks and evaluated on MDB Drums and ENST-Drums, the proposed pipeline consistently improves over a strong U-Net-based drum separation baseline in overall transcription F1. It also outperforms a representative end-to-end ADT system on kick and snare F1 under our evaluation protocol, while additionally providing separated audio stems. The ablation results show that OB gives the most stable transcription gains, whereas TB changes the trade-off between reconstruction, acoustic stem quality, and onset detection. These results suggest that generative drum demixing can serve not only as a source separation model, but also as a practical front end for interpretable drum transcription.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。