arXiv:2507.01339cs.SDcs.AI2025-07被引 3

用户哼唱或输入音谱即可分离任意乐器,灵活度远超传统四类分离。

User-guided Generative Source Separation

  • 用哼唱声波+频谱掩码双重引导,实现无固定类别的乐器分离
  • 在真实音乐中成功分离出非标准乐器,且主观评价质量高
  • 适合需要个性化提取特定旋律的音乐创作与修复场景

音乐源分离旨在从混合音频中提取单个乐器信号。现有方法多聚焦于四类分离(人声、贝斯、鼓和其他乐器),难以满足实际应用的灵活性需求。为此,我们提出GuideSep——一种基于扩散模型的无类别乐器分离方法,可突破四类限制。该模型通过两种条件输入进行控制:波形模仿条件(可通过哼唱或演奏目标旋律生成)和频谱域掩码,提供更灵活的分离引导。与依赖固定标签或声音查询的先前方法不同,本方案结合生成式框架,显著提升适用性。此外,我们设计了同架构的掩码预测基线,系统比较预测与生成方法。主客观评估表明,GuideSep不仅实现高质量分离,还支持多样化的乐器提取,验证了用户参与扩散生成过程在音乐源分离中的潜力。代码与演示页面见https://yutongwen.github.io/GuideSep/

原文摘要 · Abstract (English)

Music source separation (MSS) aims to extract individual instrument sources from their mixture. While most existing methods focus on the widely adopted four-stem separation setup (vocals, bass, drums, and other instruments), this approach lacks the flexibility needed for real-world applications. To address this, we propose GuideSep, a diffusion-based MSS model capable of instrument-agnostic separation beyond the four-stem setup. GuideSep is conditioned on multiple inputs: a waveform mimicry condition, which can be easily provided by humming or playing the target melody, and mel-spectrogram domain masks, which offer additional guidance for separation. Unlike prior approaches that relied on fixed class labels or sound queries, our conditioning scheme, coupled with the generative approach, provides greater flexibility and applicability. Additionally, we design a mask-prediction baseline using the same model architecture to systematically compare predictive and generative approaches. Our objective and subjective evaluations demonstrate that GuideSep achieves high-quality separation while enabling more versatile instrument extraction, highlighting the potential of user participation in the diffusion-based generative process for MSS. Our code and demo page are available at https://yutongwen.github.io/GuideSep/

音乐分离扩散模型用户引导生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。