用冷扩散模型解决鼓声混响去除难题,效果优于现有方法。
A Cold Diffusion Approach for Percussive Dereverberation

- 将混响建模为确定性退化过程,用冷扩散逆向生成无混响鼓声
- 在真实与合成混响数据上均超越主流基线,尤其在域外测试表现突出
- 适合音乐制作中鼓轨去混响,对瞬态信号处理能力强
当前音频去混响研究主要集中于语音,而打击乐信号因瞬态剧烈、时序密集,仍缺乏系统探索。本文提出一种冷扩散框架,用于立体声鼓声干音(混合音轨)的去混响处理,将混响视为从无混响信号逐步演变为混响信号的确定性退化过程。我们研究了两种反向过程参数化方式:直接预测(下一状态)和归一化残差(类速度)预测,并分别采用UNet与扩散Transformer作为主干网络。模型在精心筛选的包含真实与电子鼓录音的数据集上训练与评估,混响由合成与真实房间冲激响应组合生成。大量实验表明,该方法在同域与完全跨域测试集上均显著优于强基线方法,且在针对打击乐优化的信号与感知指标上表现优异。
原文摘要 · Abstract (English)
Most recent advances in audio dereverberation focus almost exclusively on speech, leaving percussive and drum signals largely unexplored despite their importance in music production. Percussive dereverberation poses distinct challenges due to sharp transients and dense temporal structure. In this work, we propose a cold diffusion framework for dereverberating stereo drum stems (downmixes), modeling reverberation as a deterministic degradation process that progressively transforms anechoic signals into reverberant ones. We investigate two reverse-process parameterizations, Direct (next-state) and a Delta-normalized residual (velocity-style) prediction, and implement the framework using both a UNet and a diffusion Transformer backbone. The models are trained and evaluated on curated datasets comprising both acoustic and electronic drum recordings, with reverberation generated using a combination of synthetic and real room impulse responses. Extensive experiments on in-domain and fully out-of-domain test sets demonstrate that the proposed method consistently outperforms strong score-based and conditional diffusion baselines, evaluated using signal-based and perceptual metrics tailored to percussive audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。