用预训练模型直接实现音乐相似性感知对齐,无需微调。
Interpretable and Perceptually-Aligned Music Similarity with Pretrained Embeddings
- 利用预训练文本-音频嵌入,结合源分离与线性优化对齐听觉偏好。
- 在ABX听觉测试中达到与微调模型相当的感知对齐效果。
- 可解释的乐器级权重,适合音乐制作人精准检索音轨素材。
感知相似性表示使音乐检索系统能判断哪些歌曲对听者而言听起来最相似。当前基于自监督度量学习的任务特定训练方法虽与人类判断有较好对齐,但因数据集有限,难以解释且泛化性差。我们发现,无需任何微调,预训练的文本-音频嵌入(CLAP 和 MuQ-MuLan)已在相似性任务上展现出可比的感知对齐效果。为超越此基线,我们提出一种新方法:通过源分离和在线性优化中使用来自听觉测试的ABX偏好数据,对预训练嵌入进行感知对齐。该模型提供可解释且可控的乐器级权重,使音乐制作者可根据混音参考曲目检索分轨循环与样本。
原文摘要 · Abstract (English)
Perceptual similarity representations enable music retrieval systems to determine which songs sound most similar to listeners. State-of-the-art approaches based on task-specific training via self-supervised metric learning show promising alignment with human judgment, but are difficult to interpret or generalize due to limited dataset availability. We show that pretrained text-audio embeddings (CLAP and MuQ-MuLan) offer comparable perceptual alignment on similarity tasks without any additional fine-tuning. To surpass this baseline, we introduce a novel method to perceptually align pretrained embeddings with source separation and linear optimization on ABX preference data from listening tests. Our model provides interpretable and controllable instrument-wise weights, allowing music producers to retrieve stem-level loops and samples based on mixed reference songs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。