arXiv:2505.16119cs.SDeess.AS2025-05中稿 · WASPAA 2025被引 9

用流匹配方法实现单通道语音分离,保证混合信号一致性。

Source Separation by Flow Matching

  • 通过流匹配学习源信号与混合信号间的微分方程映射。
  • 在真实语音混合数据上实现高保真分离,避免信号失真。
  • 适用于重叠语音分离,适合需要严格信号一致性的场景。

本文研究单通道音频源分离问题,目标是从混合信号中重建出 K 个源信号。针对该病态问题,提出 FLOSS(FLOw matching for Source Separation)方法,一种基于流匹配的受限生成模型,确保严格的混合一致性。流匹配是一种通用方法,当给定同一空间上的两个概率分布样本时,可学习一个常微分方程,在输入一个分布的样本时输出另一个分布的样本。在本工作中,我们拥有 K 个源信号的联合分布样本,以及其对应的一维混合信号分布样本。为应用流匹配,将混合信号样本通过添加人工噪声成分以匹配源信号的维度。此外,由于源信号的任意排列均产生相同混合信号,采用等变形式的流匹配,依赖于具有等变特性的神经网络架构。实验表明,该方法在重叠语音分离任务中表现优异。

原文摘要 · Abstract (English)

We consider the problem of single-channel audio source separation with the goal of reconstructing $K$ sources from their mixture. We address this ill-posed problem with FLOSS (FLOw matching for Source Separation), a constrained generation method based on flow matching, ensuring strict mixture consistency. Flow matching is a general methodology that, when given samples from two probability distributions defined on the same space, learns an ordinary differential equation to output a sample from one of the distributions when provided with a sample from the other. In our context, we have access to samples from the joint distribution of $K$ sources and so the corresponding samples from the lower-dimensional distribution of their mixture. To apply flow matching, we augment these mixture samples with artificial noise components to match the dimensionality of the $K$ source distribution. Additionally, as any permutation of the sources yields the same mixture, we adopt an equivariant formulation of flow matching which relies on a neural network architecture that is equivariant by design. We demonstrate the performance of the method for the separation of overlapping speech.

语音分离流匹配生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。