构建吉他二重奏数据集,提升相似音色乐器分离效果
Classical Guitar Duet Separation using GuitarDuets -- a Dataset of Real and Synthesized Guitar Recordings
- 用真实与合成录音构建吉他二重奏数据集
- 结合音符标注可提升分离性能,合成数据更有效
- 适合音乐信号处理与音频分离研究者
音乐源分离(MSS)近年多聚焦于多音色场景,现有模型针对不同乐器设计,忽视了音色相近乐器的分离挑战。本文聚焦单调色MSS,以古典吉他二重奏为对象。我们提出GuitarDuets数据集,包含约三小时真实与合成吉他二重奏录音,以及合成部分的音符级标注。通过适配前沿的Demucs架构进行跨数据集评估,并提出联合排列不变音符转录与分离框架,利用音符事件作为辅助信息。结果表明,同时使用真实与合成数据集,在独立测试集上优于仅用单一数据;尽管真实音符标签显著提升性能,预测音符估计仅带来微弱改善。最后讨论了SDR与SI-SDR等常用指标在单调色分离中的表现。
原文摘要 · Abstract (English)
Recent advancements in music source separation (MSS) have focused in the multi-timbral case, with existing architectures tailored for the separation of distinct instruments, overlooking thus the challenge of separating instruments with similar timbral characteristics. Addressing this gap, our work focuses on monotimbral MSS, specifically within the context of classical guitar duets. To this end, we introduce the GuitarDuets dataset, featuring a combined total of approximately three hours of real and synthesized classical guitar duet recordings, as well as note-level annotations of the synthesized duets. We perform an extensive cross-dataset evaluation by adapting Demucs, a state-of-the-art MSS architecture, to monotimbral source separation. Furthermore, we develop a joint permutation-invariant transcription and separation framework, to exploit note event predictions as auxiliary information. Our results indicate that utilizing both the real and synthesized subsets of GuitarDuets leads to improved separation performance in an independently recorded test set compared to utilizing solely one subset. We also find that while the availability of ground-truth note labels greatly helps the performance of the separation network, the predicted note estimates result only in marginal improvement. Finally, we discuss the behavior of commonly utilized metrics, such as SDR and SI-SDR, in the context of monotimbral MSS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。