arXiv:2608.22111cs.SD2026-08

用生成式方法根据文字描述分离音频,效果优于传统方法。

FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation

论文配图:FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation
图 1 · 摘自论文原文
  • 通过扩散模型从噪声中生成目标音源表示,不依赖掩码预测。
  • 在复杂重叠场景下,分离效果显著提升,超越现有最佳方法。
  • 适合需要高精度音频分离的语音处理、影视制作等场景。

语言查询音频源分离(LASS)旨在根据自然语言描述从混音中提取目标声源,提供灵活可扩展的音频分离接口。然而,现有大多数LASS方法依赖判别性掩码模型,直接从输入混音估计掩码,常导致目标声音过度抑制或未能充分分离,尤其在多个声事件高度重叠的复杂声景中表现不佳。本文提出FlowSep2,一种文本条件的流匹配生成模型用于LASS。不同于直接预测分离掩码,FlowSep2在潜在空间中从高斯噪声生成目标声源表示,同时以混合信号表示和文本查询为条件。我们采用带有扩散变压器主干的修正流匹配方法,并引入自监督流匹配机制Self-Flow,使潜在表示在生成目标下具备语义结构,从而增强按文本查询分离目标声源的能力。在多个LASS基准测试中,FlowSep2达到当前最优性能,在包含重叠声事件的挑战性场景中表现出更优的分离效果。

原文摘要 · Abstract (English)

Language-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable interface for audio source separation. However, most existing LASS methods rely on discriminative, mask-based models, which estimate masks from the input mixture. These methods often over-suppress target sounds or fail to fully separate them, especially when multiple sound events strongly overlap in complex acoustic scenes. In this work, we propose FlowSep2, a text-conditioned flow-matching generative model for LASS. Instead of directly predicting a separation mask, FlowSep2 learns to generate the target source representation from Gaussian noise in a latent space, conditioned on both the mixture representation and the text query. Specifically, we employ rectified flow matching with a Diffusion Transformer backbone. We further incorporate Self-Flow, a self-supervised flow-matching paradigm, into our LASS framework. By encouraging semantically structured latent representations under the generative objective, Self-Flow improves the model's ability to separate target sources according to text queries. Experiments on multiple LASS benchmarks show that FlowSep2 achieves state-of-the-art performance and demonstrates enhanced sound separation results in challenging scenarios with overlapping sound events.

音频分离生成模型文本控制流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。