arXiv:2501.06959cs.SDcs.DL2025-01中稿 · the 25th Internati…被引 5

首个高质量卡纳提克音乐多模态数据集,助力非西方音乐源分离。

Sanidha: A Studio Quality Multi-Modal Dataset for Carnatic Music

  • 构建首个开源卡纳提克音乐多轨录音数据集,音轨间重叠极少
  • 在新数据集上微调Spleeter,信噪比提升显著优于旧数据集
  • 配套高清视频,适合音乐信息检索与跨文化音频研究者

音乐源分离旨在将一段音乐拆分为独立声源(如人声、打击乐、旋律乐器等),该任务无简单数学解法,需依赖大规模孤立音轨数据集进行深度学习训练。现有主流数据集多基于商业西方音乐,限制了模型在非西方音乐如卡纳提克音乐中的应用。卡纳提克音乐为现场传统,多轨录音常含声源重叠与串扰,对Spleeter和Hybrid Demucs等商用分离模型构成挑战。本文提出‘Sanidha’,首个面向卡纳提克音乐的开源高质量多模态数据集,提供低重叠、无串扰的多轨音频及艺术家表演高清视频。我们还在该数据集上微调了常用源分离模型Spleeter,相较在已有卡纳提克数据集上的微调,其信噪比(SDR)性能显著提升。经听觉评估,新模型输出质量更优。

原文摘要 · Abstract (English)

Music source separation demixes a piece of music into its individual sound sources (vocals, percussion, melodic instruments, etc.), a task with no simple mathematical solution. It requires deep learning methods involving training on large datasets of isolated music stems. The most commonly available datasets are made from commercial Western music, limiting the models' applications to non-Western genres like Carnatic music. Carnatic music is a live tradition, with the available multi-track recordings containing overlapping sounds and bleeds between the sources. This poses a challenge to commercially available source separation models like Spleeter and Hybrid Demucs. In this work, we introduce 'Sanidha', the first open-source novel dataset for Carnatic music, offering studio-quality, multi-track recordings with minimal to no overlap or bleed. Along with the audio files, we provide high-definition videos of the artists' performances. Additionally, we fine-tuned Spleeter, one of the most commonly used source separation models, on our dataset and observed improved SDR performance compared to fine-tuning on a pre-existing Carnatic multi-track dataset. The outputs of the fine-tuned model with 'Sanidha' are evaluated through a listening study.

音乐分离卡纳提克多模态数据集音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。