arXiv:2512.02432cs.SD2025-12

用户可交互式调整歌声分离模型,实时提升音轨处理效果。

Continual Learning for Singing Voice Separation with Human in the Loop Adaptation

  • 引入人机协同持续学习框架,用户通过标记错误区域反馈改进模型。
  • 在同数据集和跨数据集场景下,分离性能显著优于基础模型。
  • 适合音乐制作、音频编辑等需个性化调整的实用场景。

基于深度学习的歌声分离方法近期表现优异,但多数未考虑用户交互以提升性能。这在真实场景中尤为关键,因实际音乐曲目可能与训练数据在风格和伴奏乐器上存在差异。本文提出一种基于U-Net的交互式持续学习框架,允许用户通过标记提取出的歌声中本应为静音的误判区域,对模型进行微调。设计了两种持续学习算法,在同数据集和跨数据集设置下均验证了其相比基础模型的性能提升。

原文摘要 · Abstract (English)

Deep learning-based works for singing voice separation have performed exceptionally well in the recent past. However, most of these works do not focus on allowing users to interact with the model to improve performance. This can be crucial when deploying the model in real-world scenarios where music tracks can vary from the original training data in both genre and instruments. In this paper, we present a deep learning-based interactive continual learning framework for singing voice separation that allows users to fine-tune the vocal separation model to conform it to new target songs. We use a U-Net-based base model architecture that produces a mask for separating vocals from the spectrogram, followed by a human-in-the-loop task where the user provides feedback by marking a few false positives, i.e., regions in the extracted vocals that should have been silence. We propose two continual learning algorithms. Experiments substantiate the improvement in singing voice separation performance by the proposed algorithms over the base model in intra-dataset and inter-dataset settings.

歌声分离持续学习人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。