用改进的扩散模型实现高效人声分离,性能接近顶尖水平
Diff-VS: Efficient Audio-Aware Diffusion U-Net for Vocals Separation
- 基于音乐设计优化的U-Net结构处理时频谱图
- 在客观指标上媲美判别式基线,主观感知质量领先
- 适合对音频生成与分离感兴趣的开发者
尽管扩散模型以生成任务著称,也已成功应用于多种任务,包括音频源分离。然而,当前基于生成的方法在标准客观指标上表现不佳。本文提出一种基于阐明扩散模型(EDM)框架的新颖生成式人声分离模型,处理复杂的短时傅里叶变换谱图,采用受音乐启发设计的改进U-Net架构。该方法在客观指标上达到判别式基线水平,且通过代理主观评估,感知质量接近当前最优系统。这些结果希望推动生成方法在音乐源分离中的更广泛应用。
原文摘要 · Abstract (English)
While diffusion models are best known for their performance in generative tasks, they have also been successfully applied to many other tasks, including audio source separation. However, current generative approaches to music source separation often underperform on standard objective metrics. In this paper, we address this issue by introducing a novel generative vocal separation model based on the Elucidated Diffusion Model (EDM) framework. Our model processes complex short-time Fourier transform spectrograms and employs an improved U-Net architecture based on music-informed design choices. Our approach matches discriminative baselines on objective metrics and achieves perceptual quality comparable to state-of-the-art systems, as assessed by proxy subjective metrics. We hope these results encourage broader exploration of generative methods for music source separation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。