让音乐相似度模型区分旋律、节奏和音色,提升可解释性与控制力。
MERIT: Learning Disentangled Music Representations for Audio Similarity

- 通过条件生成与分离音轨,训练时只改变单一音乐维度。
- 各维度表示头在对应特征上响应强,其他维度接近随机水平。
- 适用于需要精准音乐检索与分析的场景,如智能编曲或跨风格匹配。
当前音乐相似度模型通常输出单一综合分数,将旋律、节奏、音色等音乐维度混合,限制了用户控制力和可解释性,难以实现细粒度查询。本文提出MERIT框架,学习针对这三大核心维度的解耦表示。为克服真实音频中缺乏独立变化的问题,采用新型训练策略:利用条件音频生成与源分离音轨,强化训练数据中单一因素的变化。评估表明,模型在因子层面具有显著解耦性:每个表示头对预期感知维度响应强烈,而在其他维度上接近随机水平,这一特性在合成训练域和独立真实音频上均成立。
原文摘要 · Abstract (English)
Current music similarity models typically compute a single, monolithic score, entangling distinct musical dimensions like melody, rhythm, and timbre. This limits user control and interpretability, making it impossible to execute nuanced queries. We introduce MERIT, a framework for learning disentangled, factor-specific music representations tailored to these three core dimensions. To overcome the lack of isolated musical variations in real-world audio, we use a novel training strategy that uses conditional audio generation and source-separated stems to strongly encourage single-factor variation in training data. Our evaluations demonstrate strong factor-wise disentanglement. Each head responds strongly to its intended perceptual dimension while remaining near chance on the others, a representational property that holds across both the synthetic training domain and independent real-world audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。