无需训练即可分离任意音频源,靠文本提示引导扩散模型完成
ZeroSep: Separate Anything in Audio with Zero Training
- 用文本条件控制扩散模型的反向去噪过程分离音源
- 零样本下在多个基准上超越有监督方法性能
- 适合需要快速部署、应对未知音源场景的开发者
音频源分离是机器理解复杂声学环境的基础,支撑众多音频应用。现有监督深度学习方法受限于大量特定任务标注数据,难以泛化到真实世界中无限多样的声学场景。受生成式基础模型成功启发,我们探索预训练的文本引导音频扩散模型是否能克服这些限制。令人惊讶的是,仅通过正确配置的预训练文本引导音频扩散模型,即可实现零样本源分离。我们的方法ZeroSep将混合音频反推至扩散模型隐空间,利用文本条件引导去噪过程以恢复单个音源。无需任何特定任务训练或微调,ZeroSep将生成式扩散模型复用于判别性分离任务,并通过丰富的文本先验自然支持开放集场景。该方法兼容多种预训练文本引导音频扩散主干网络,在多个分离基准上表现优异,甚至超越部分有监督方法。
原文摘要 · Abstract (English)
Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive, task-specific labeled data and struggle to generalize to the immense variability and open-set nature of real-world acoustic scenes. Inspired by the success of generative foundation models, we investigate whether pre-trained text-guided audio diffusion models can overcome these limitations. We make a surprising discovery: zero-shot source separation can be achieved purely through a pre-trained text-guided audio diffusion model under the right configuration. Our method, named ZeroSep, works by inverting the mixed audio into the diffusion model's latent space and then using text conditioning to guide the denoising process to recover individual sources. Without any task-specific training or fine-tuning, ZeroSep repurposes the generative diffusion model for a discriminative separation task and inherently supports open-set scenarios through its rich textual priors. ZeroSep is compatible with a variety of pre-trained text-guided audio diffusion backbones and delivers strong separation performance on multiple separation benchmarks, surpassing even supervised methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。