arXiv:2509.13285cs.SDcs.AI2025-09被引 1

用对比学习提升乐器音色检索精度,支持单/多乐器混合场景。

Contrastive timbre representations for musical instrument and synthesizer retrieval

  • 采用对比学习框架,生成真实虚拟乐器的正负样本对。
  • 三乐器混合输入下,顶1准确率达81.7%,顶5达95.7%。
  • 适用于数字音乐制作中快速定位特定音色,适合音频检索研究者。

在数字音乐制作中,从音频混音中高效检索特定乐器音色仍具挑战。本文提出一种用于乐器检索的对比学习框架,支持使用单一模型直接查询乐器数据库,涵盖单乐器与多乐器声音。我们设计了生成真实虚拟乐器(如采样器、合成器)正负样本对的技术,克服了传统音频数据增强方法的局限性。首次实验在包含3,884种乐器的数据集上进行,以单乐器音频为输入,对比方法性能与基于分类预训练的现有工作相当。第二次实验针对多乐器混合音频输入的检索任务,所提框架表现更优,在三乐器混合情况下达到81.7%的顶1准确率和95.7%的顶5准确率。

原文摘要 · Abstract (English)

Efficiently retrieving specific instrument timbres from audio mixtures remains a challenge in digital music production. This paper introduces a contrastive learning framework for musical instrument retrieval, enabling direct querying of instrument databases using a single model for both single- and multi-instrument sounds. We propose techniques to generate realistic positive/negative pairs of sounds for virtual musical instruments, such as samplers and synthesizers, addressing limitations in common audio data augmentation methods. The first experiment focuses on instrument retrieval from a dataset of 3,884 instruments, using single-instrument audio as input. Contrastive approaches are competitive with previous works based on classification pre-training. The second experiment considers multi-instrument retrieval with a mixture of instruments as audio input. In this case, the proposed contrastive framework outperforms related works, achieving 81.7\% top-1 and 95.7\% top-5 accuracies for three-instrument mixtures.

音色检索对比学习音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。