arXiv:2503.17281cs.SDeess.AS2025-03被引 3

用单一模型分离乐器特征,实现更准的音乐相似性匹配

Learning Separated Representations for Instrument-based Music Similarity

  • 单模型输入混音信号,通过条件相似网络分离乐器子空间
  • 在低准确率乐器上表现优于独立网络,且子空间保留乐器特性
  • 可精准聚焦特定乐器音色,用户接受度高

灵活的推荐与检索系统需支持对乐曲多个局部元素的相似性判断。现有基于多网络和单独乐器信号的方法虽有效,但使用纯净乐器信号作为查询不适用于实际检索系统,且分离后的信号因失真导致精度下降。本文提出一种基于单网络的乐器部分音乐相似性学习方法,直接以混合信号为输入。设计了一个包含各乐器分离子空间的统一相似性嵌入空间,由条件相似网络通过带掩码的三元组损失训练。实验表明:(1)在低准确率乐器上,该方法获得的嵌入表示比使用分离信号的独立网络更准确;(2)每个子嵌入空间能有效保留对应乐器的特征;(3)用户评估显示,该方法在聚焦特定乐器音色时具有较高接受度。

原文摘要 · Abstract (English)

A flexible recommendation and retrieval system requires music similarity in terms of multiple partial elements of musical pieces to allow users to select the element they want to focus on. A method for music similarity learning using multiple networks with individual instrumental signals is effective but faces the problem that using each clean instrumental signal as a query is impractical for retrieval systems and using separated instrumental signals reduces accuracy owing to artifacts. In this paper, we present instrumental-part-based music similarity learning with a single network that takes mixed signals as input instead of individual instrumental signals. Specifically, we designed a single similarity embedding space with separated subspaces for each instrument, extracted by Conditional Similarity Networks, which are trained using the triplet loss with masks. Experimental results showed that (1) the proposed method can obtain more accurate embedding representation than using individual networks using separated signals as input in the evaluation of an instrument that had low accuracy, (2) each sub-embedding space can hold the characteristics of the corresponding instrument, and (3) the selection of similar musical pieces focusing on each instrumental sound by the proposed method can obtain human acceptance, especially when focusing on timbre.

音乐相似性乐器分离嵌入空间条件网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。