用少步生成提升人声伴奏分离效果,兼顾速度与质量。
Decoupled Latent Flow Matching for Few-Step Joint Vocal-Accompaniment Separation

- 分离编码器与速度解码器解耦,分别处理语义和声学信息。
- 通过对抗后训练实现少步采样,提升生成效率与音质。
- 适合需要快速高质量音乐分离的应用场景。
生成建模为建模混合信号条件下的源分布提供了灵活性,但迭代扩散和流匹配模型在处理长音乐信号时成本较高。本文通过潜在空间流匹配研究联合人声与伴奏分离,利用预训练变分自编码器(VAE)将混合信号与源信号映射到紧凑的潜在空间,并通过流匹配模型联合生成人声与伴奏的潜在表示。所提框架通过分离编码器与速度解码器解耦语义分离与声学速度预测。为降低采样成本,进一步采用受Flow2GAN启发的潜在对抗后训练,实现少步生成。实验表明,在减少采样预算的情况下,潜在对抗精炼可提升感知质量和分离指标。
原文摘要 · Abstract (English)
Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。