融合多模型优势,提升音乐分离精度与多样性
An Ensemble Approach to Music Source Separation: A Comparative Analysis of Conventional and Hierarchical Stem Separation
- 采用集成方法结合多个先进模型,互补优势提升性能
- 在VDB三类主音轨上表现稳定,次级子音轨分离效果显著
- 适合追求高精度音轨分离的研究者与音频工程应用
音乐源分离(MSS)旨在从混合音频中提取独立声源(即音轨)。本文提出一种集成方法,融合多个前沿模型架构,在传统人声、鼓组和贝斯(VDB)音轨上实现更优分离效果,并拓展至二级层次分离,如底鼓、军鼓、主唱和伴唱等子音轨。通过综合信噪比(SNR)与信号失真比(SDR)的调和平均值进行音轨评估,避免极端值干扰,均衡两指标权重。实验表明,该方法在各类主流音轨上均保持高水平表现,同时揭示了流派与乐器配置对模型性能的影响。尽管二级分离仍有提升空间,但实现子音轨精细分离已具里程碑意义。研究为后续扩展模型能力至吉他、钢琴等细分音轨提供新方向。
原文摘要 · Abstract (English)
Music source separation (MSS) is a task that involves isolating individual sound sources, or stems, from mixed audio signals. This paper presents an ensemble approach to MSS, combining several state-of-the-art architectures to achieve superior separation performance across traditional Vocal, Drum, and Bass (VDB) stems, as well as expanding into second-level hierarchical separation for sub-stems like kick, snare, lead vocals, and background vocals. Our method addresses the limitations of relying on a single model by utilising the complementary strengths of various models, leading to more balanced results across stems. For stem selection, we used the harmonic mean of Signal-to-Noise Ratio (SNR) and Signal-to-Distortion Ratio (SDR), ensuring that extreme values do not skew the results and that both metrics are weighted effectively. In addition to consistently high performance across the VDB stems, we also explored second-level hierarchical separation, revealing important insights into the complexities of MSS and how factors like genre and instrumentation can influence model performance. While the second-level separation results show room for improvement, the ability to isolate sub-stems marks a significant advancement. Our findings pave the way for further research in MSS, particularly in expanding model capabilities beyond VDB and improving niche stem separations such as guitar and piano.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。