arXiv:2605.27346cs.SD2026-05

让音乐相似度模型区分旋律、节奏和音色,提升可解释性与控制力。

MERIT: Learning Disentangled Music Representations for Audio Similarity

论文配图:MERIT: Learning Disentangled Music Representations for Audio Similarity
图 1 · 摘自论文原文
  • 通过条件生成与分离音轨,训练时只改变单一音乐维度。
  • 各维度表示头在对应特征上响应强,其他维度接近随机水平。
  • 适用于需要精准音乐检索与分析的场景,如智能编曲或跨风格匹配。

当前音乐相似度模型通常输出单一综合分数,将旋律、节奏、音色等音乐维度混合,限制了用户控制力和可解释性,难以实现细粒度查询。本文提出MERIT框架,学习针对这三大核心维度的解耦表示。为克服真实音频中缺乏独立变化的问题,采用新型训练策略:利用条件音频生成与源分离音轨,强化训练数据中单一因素的变化。评估表明,模型在因子层面具有显著解耦性:每个表示头对预期感知维度响应强烈,而在其他维度上接近随机水平,这一特性在合成训练域和独立真实音频上均成立。

原文摘要 · Abstract (English)

Current music similarity models typically compute a single, monolithic score, entangling distinct musical dimensions like melody, rhythm, and timbre. This limits user control and interpretability, making it impossible to execute nuanced queries. We introduce MERIT, a framework for learning disentangled, factor-specific music representations tailored to these three core dimensions. To overcome the lack of isolated musical variations in real-world audio, we use a novel training strategy that uses conditional audio generation and source-separated stems to strongly encourage single-factor variation in training data. Our evaluations demonstrate strong factor-wise disentanglement. Each head responds strongly to its intended perceptual dimension while remaining near chance on the others, a representational property that holds across both the synthetic training domain and independent real-world audio.

音乐表示解耦表征音频相似度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。