arXiv:2608.14516eess.AScs.SD2026-08中稿 · IWAENC 2026

通过歌手音色嵌入,精准分离多人演唱音乐中的目标人声。

Singer-Informed Vocal Source Separation for Multi-Singer Music Mixtures

论文配图:Singer-Informed Vocal Source Separation for Multi-Singer Music Mixtures
图 1 · 摘自论文原文
  • 用短段录音生成歌手嵌入,通过特征拼接或FiLM调制引导分离。
  • 多歌手场景下目标人声SI-SDR提升至5.58 dB(原0.33 dB)。
  • 适合需要精准分离特定歌手的音乐制作与语音处理应用。

现有音乐源分离系统通常只提取单一人声,无法区分多人演唱。本文研究面向多歌手混合音频的歌手知情人声分离方法。框架引入目标歌手的短段录音作为参考,生成学习得到的歌手嵌入,通过特征拼接或特征式线性调制(FiLM)融入模型,使模型聚焦于目标歌手并抑制干扰。基于DAMP-VSEP构建了经过质量筛选和非重叠录段处理的二重唱数据集。在独唱与对唱设置下的实验表明,尽管基线模型在单歌手混合中表现良好,但本文方法在多歌手场景下显著提升了目标歌手的提取效果,使目标歌手的SI-SDR从0.33 dB提升至5.58 dB。Fréchet Audio Distance(FAD)结果进一步显示,分离音频的听觉质量更优,且与目标音频分布匹配度更高。代码与模型检查点已公开于https://github.com/jocelynxu01/singer-separation-paper。

原文摘要 · Abstract (English)

Music source separation systems typically extract a single vocal track and do not distinguish between multiple singers. We study singer-informed vocal source separation for multi-singer mixtures. Our framework introduces a short enrollment recording of a target singer to guide separation through a learned embedding. The singer embedding is incorporated using feature concatenation or feature-wise linear modulation (FiLM), enabling the model to focus on the target singer while suppressing interference. We construct a duet dataset based on DAMP-VSEP with quality filtering and non-overlapping enrollment segments. Experiments on solo and duet settings show that while baseline models perform well for single-singer mixtures, the proposed method improves target-singer extraction in multi-singer cases, increasing target-singer SI-SDR from 0.33 dB to 5.58 dB. Fréchet Audio Distance (FAD) further shows improved perceptual quality and better alignment with target audio distributions. Code and checkpoints are available at https://github.com/jocelynxu01/singer-separation-paper.

人声分离多歌手音频处理嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。