arXiv:2607.23395cs.SDcs.LG2026-07

统一音乐分离训练与评估流程,提升模型开发效率。

Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

  • 用配置文件统一管理模型训练、验证和推理流程。
  • 支持多种模型架构与增强技术,显著提升分离效果。
  • 适合音频处理研究者快速实验与复现结果。

音乐源分离(MSS)旨在从多声部混合音频中还原出单独的声音成分(音轨),广泛应用于卡拉OK、混音、音频修复和内容创作等场景。分离质量受整个流程中多项工程决策影响:模型选择、训练数据准备与增强、损失函数与评估指标设计、训练配置、验证策略及后处理。本文提出MSST(Music-Source-Separation-Training)——一个面向MSS任务的通用开源框架,通过单一配置驱动接口,统一支持多种现代分离模型家族的训练、验证与推理。该框架支持多样化的模型架构、数据预处理与增强方法、多种损失函数与评估指标,便于快速迭代与消融实验。同时集成滑动窗口推理、交叉淡入、测试时增强、模型集成及低秩适应(LORA)微调等实用技术,实证显示可有效提升分离性能。通过将上述组件整合为可复现、基于YAML配置的框架,MSST降低了系统性实验门槛,实现从想法到可验证结果的快速转化。

原文摘要 · Abstract (English)

Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production. The separation quality depends on engineering decisions across the entire pipeline: model choice, training data preparation and augmentation, loss function and metrics choice, training configuration, validation, and post-processing. This paper presents MSST (Music-Source-Separation-Training) - a universal open-source framework for MSS tasks, which unifies training, validation, and inference for a broad range of modern demixing model families under a single, configuration-driven interface. The framework supports various model architectures, data preprocessing and augmentations, multiple loss functions and evaluation metrics, which helps with fast iterations and ablation studies. Additionally, the framework supports a range of practical techniques that improve separation quality, such as sliding-window inference with cross-fading, test-time augmentation, model ensembling, and fine-tuning via Low-Rank Adaptation (LORA). Our ablation studies demonstrate improvements of MSS using the above techniques. By consolidating these components into a reproducible, YAML-configurable framework, MSST lowers the barrier to systematic experimentation and enables rapid iteration from idea to verifiable result.

音乐分离深度学习音频处理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。