MixIT可有效用于音乐源分离预训练,提升模型性能。
Is MixIT Really Unsuitable for Correlated Sources? Exploring MixIT for Unsupervised Pre-training in Music Source Separation
- 用MixIT在无标签音乐数据上预训练模型
- 在MUSDB18上微调后性能优于从零训练
- 适合缺乏标注数据的音乐分离场景
在音乐源分离(MSS)中,获取孤立音轨成本高昂,因此在无标签数据上进行预训练成为可行方案。尽管如混合不变训练(MixIT)这类无需源假设的无监督学习方法在通用声音分离中已有探索,但在MSS领域却因隐含的源独立性假设而被忽视。我们提出,其应用困难源于MSS本身的病态性:音轨定义依赖具体应用,模型缺乏应分离或不应分离的先验知识,而非高源相关性所致。尽管MixIT不依赖源模型且对模糊性敏感,初步实验表明其仍能部分分离乐器,显示出潜在价值。受此启发,本研究探索基于MixIT的MSS预训练策略。首先使用Free Music Archive中的真实世界无标签数据进行MixIT预训练,再在带标签的MUSDB18上微调。采用当前先进的band-split TF-Locoformer模型,结果表明,该预训练策略显著优于从零开始训练。
原文摘要 · Abstract (English)
In music source separation (MSS), obtaining isolated sources or stems is highly costly, making pre-training on unlabeled data a promising approach. Although source-agnostic unsupervised learning like mixture-invariant training (MixIT) has been explored in general sound separation, they have been largely overlooked in MSS due to its implicit assumption of source independence. We hypothesize, however, that the difficulty of applying MixIT to MSS arises from the ill-posed nature of MSS itself, where stem definitions are application-dependent and models lack explicit knowledge of what should or should not be separated, rather than from high inter-source correlation. While MixIT does not assume any source model and struggles with such ambiguities, our preliminary experiments show that it can still separate instruments to some extent, suggesting its potential for unsupervised pre-training. Motivated by these insights, this study investigates MixIT-based pre-training for MSS. We first pre-train a model on in-the-wild, unlabeled data from the Free Music Archive using MixIT, and then fine-tune it on MUSDB18 with supervision. Using the band-split TF-Locoformer, one of the state-of-the-art MSS models, we demonstrate that MixIT-based pre-training improves the performance over training from scratch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。