arXiv:2510.06785eess.AScs.SD2025-10被引 3

轻量级音乐分离模型,参数少13倍仍保持高精度。

Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation

  • 采用频段分割U-Net架构,优化计算资源分配。
  • 在MUSDB-HQ上达到与大模型相当的音源分离效果。
  • 适合移动端或边缘设备部署,支持数据扩展。

近年来,音乐源分离领域取得了显著进展,双路径建模、频段分割模块和Transformer层等架构已实现良好性能。然而,这些模型通常参数量庞大,在计算资源受限的设备上训练和应用面临挑战。尽管已有轻量级模型提出,但其性能普遍低于大型模型。本文受近期进展启发,通过精心设计,使轻量级模型在分离四类音乐音轨时达到与参数量高达13倍的模型相当的信噪比(SDR)表现。所提出的Moises-Light模型在MUSDB-HQ基准数据集上表现优异,并在使用额外训练数据MoisesDB时展现出良好的可扩展性。

原文摘要 · Abstract (English)

In recent years, significant advances have been made in music source separation, with model architectures such as dual-path modeling, band-split modules, or transformer layers achieving comparably good results. However, these models often contain a significant number of parameters, posing challenges to devices with limited computational resources in terms of training and practical application. While some lightweight models have been introduced, they generally perform worse compared to their larger counterparts. In this paper, we take inspiration from these recent advances to improve a lightweight model. We demonstrate that with careful design, a lightweight model can achieve comparable SDRs to models with up to 13 times more parameters. Our proposed model, Moises-Light, achieves competitive results in separating four musical stems on the MUSDB-HQ benchmark dataset. The proposed model also demonstrates competitive scalability when using MoisesDB as additional training data.

音乐分离轻量化频段分割U-Net

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。