arXiv:2409.10995eess.AScs.LG2024-09被引 9

构建合成管弦乐分离数据集,解决真实多轨数据稀缺问题

SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation

  • 用仿真技术生成高保真、多样化的管弦乐多轨数据
  • 在合成与真实场景下验证基线模型性能,提升分离效果
  • 适合音乐信号处理与音频生成研究者使用

近年来,音乐源分离技术取得显著进展,尤其在人声、鼓点和贝斯的分离上。这得益于针对特定音源的大规模多轨数据集的建立与应用。然而,管弦乐中相似音色音源的分离仍缺乏充分探索,主要受限于高质量、无串扰的多轨数据集稀缺。本文提出一种新型多轨数据集SynthSOD,采用仿真技术生成具有高保真音色(使用优质音色库)、音乐合理性、以及多样化动态、自然节奏变化、风格和条件的训练数据。此外,我们展示了在该合成数据集上训练的主流音乐分离模型在知名EnsembleSet上的表现,并评估其在合成与真实场景下的性能,验证了数据集的有效性。

原文摘要 · Abstract (English)

Recent advancements in music source separation have significantly progressed, particularly in isolating vocals, drums, and bass elements from mixed tracks. These developments owe much to the creation and use of large-scale, multitrack datasets dedicated to these specific components. However, the challenge of extracting similarly sounding sources from orchestra recordings has not been extensively explored, largely due to a scarcity of comprehensive and clean (i.e bleed-free) multitrack datasets. In this paper, we introduce a novel multitrack dataset called SynthSOD, developed using a set of simulation techniques to create a realistic (i.e. using high-quality soundfonts), musically motivated, and heterogeneous training set comprising different dynamics, natural tempo changes, styles, and conditions. Moreover, we demonstrate the application of a widely used baseline music separation model trained on our synthesized dataset w.r.t to the well-known EnsembleSet, and evaluate its performance under both synthetic and real-world conditions.

音乐分离合成数据管弦乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。