无需干净样本,直接从混合信号中分离音源,效果超越现有无监督方法。
SURF: Separation via Unsupervised Remixing Flow

- 通过教师-学生架构与重混技术,从混合信号中无监督训练流模型。
- 在音频和图像任务上达到新最优,分离准确率显著领先现有方法。
- 适合缺乏标注数据的场景,如真实环境音源分离或跨域应用。
单通道源分离的目标是从混合信号中重建出 K 个原始信号。在拥有大量干净源数据的监督设置下,生成式扩散模型和基于流的先验模型已成功解决这一困难且病态的问题。然而,干净源样本通常难以获取,且即使可用,监督模型也易受领域偏移影响。为弥合此差距,我们提出无监督重混流分离方法(SURF),一种直接从观测混合信号中学习的无监督流匹配方法。该方法结合了最先进的监督流匹配与基于回归的自监督技术。总体而言,从教师模型出发,通过“重混”步骤,利用教师估计值来引导学生流模型的学习。本文还揭示了该方法所优化目标的本质,并建立与唤醒-睡眠算法的新关联。在图像与音频基准上的实证评估表明,SURF 建立了新的最优性能,显著优于现有无监督方法。示例见演示页面:https://google.github.io/df-conformer/surf/
原文摘要 · Abstract (English)
The goal of single-channel source separation is to reconstruct $K$ sources given their mixture. In supervised settings where vast amounts of clean source data are available, this challenging, ill-posed problem has been addressed successfully by generative diffusion and flow-based prior models. However, access to such clean source samples is often limited, and even when available, supervised models are vulnerable to domain shifts. To bridge this gap, we present Separation via Unsupervised Remixing Flow (SURF), an unsupervised flow matching approach for source separation that learns directly from observed mixtures. This method relies on a novel combination of state-of-the-art supervised flow matching and regression-based self-supervised techniques. At a high level, starting from a teacher model, we utilize a "remixing" step to bootstrap the learning of a student flow model from the teacher's estimates. We provide insights into the objectives optimized by this approach and draw a novel connection to the Wake-Sleep algorithm. Empirical evaluations on image and audio benchmarks demonstrate that SURF establishes a new state-of-the-art, significantly outperforming existing unsupervised methods. See our demo page for examples. https://google.github.io/df-conformer/surf/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。