用生成式方法根据文字描述分离音频,效果优于传统方法。
FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation

- 通过扩散模型从噪声中生成目标音源表示,不依赖掩码预测。
- 在复杂重叠场景下,分离效果显著提升,超越现有最佳方法。
- 适合需要高精度音频分离的语音处理、影视制作等场景。
语言查询音频源分离(LASS)旨在根据自然语言描述从混音中提取目标声源,提供灵活可扩展的音频分离接口。然而,现有大多数LASS方法依赖判别性掩码模型,直接从输入混音估计掩码,常导致目标声音过度抑制或未能充分分离,尤其在多个声事件高度重叠的复杂声景中表现不佳。本文提出FlowSep2,一种文本条件的流匹配生成模型用于LASS。不同于直接预测分离掩码,FlowSep2在潜在空间中从高斯噪声生成目标声源表示,同时以混合信号表示和文本查询为条件。我们采用带有扩散变压器主干的修正流匹配方法,并引入自监督流匹配机制Self-Flow,使潜在表示在生成目标下具备语义结构,从而增强按文本查询分离目标声源的能力。在多个LASS基准测试中,FlowSep2达到当前最优性能,在包含重叠声事件的挑战性场景中表现出更优的分离效果。
原文摘要 · Abstract (English)
Language-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable interface for audio source separation. However, most existing LASS methods rely on discriminative, mask-based models, which estimate masks from the input mixture. These methods often over-suppress target sounds or fail to fully separate them, especially when multiple sound events strongly overlap in complex acoustic scenes. In this work, we propose FlowSep2, a text-conditioned flow-matching generative model for LASS. Instead of directly predicting a separation mask, FlowSep2 learns to generate the target source representation from Gaussian noise in a latent space, conditioned on both the mixture representation and the text query. Specifically, we employ rectified flow matching with a Diffusion Transformer backbone. We further incorporate Self-Flow, a self-supervised flow-matching paradigm, into our LASS framework. By encouraging semantically structured latent representations under the generative objective, Self-Flow improves the model's ability to separate target sources according to text queries. Experiments on multiple LASS benchmarks show that FlowSep2 achieves state-of-the-art performance and demonstrates enhanced sound separation results in challenging scenarios with overlapping sound events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。