用预训练音乐嵌入识别爵士标准曲,提升跨演奏版本检索效果
Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition

- 对比自建模型与预训练音乐嵌入在跨演奏识别中的表现
- 预训练嵌入在top-k检索中更优,但受演奏者身份影响
- 轻量级对比投影可部分缓解演奏者偏差,适合音乐检索研究者
从音频中识别爵士标准曲是一项具有挑战性的曲目级音乐检索任务:同一首标准曲的不同演奏版本在速度、调性、编排、乐器配置、即兴内容甚至主旋律是否存在方面差异显著。本文使用专为跨演奏版本识别设计的爵士三重奏数据库子集进行研究。对比了从头训练的谐音卷积神经网络基线模型与近期音乐理解基础模型提供的冻结预训练音乐表示,采用监督探测和最近邻检索两种方法。结果表明,从头训练的频谱模型对训练演奏版本过拟合严重,而预训练嵌入在top-$k$检索中表现更好,但对演奏者身份敏感,可通过轻量级对比投影部分缓解。研究结果推动将爵士标准曲识别作为音乐表示模型的有效压力测试,并为基于检索的标准曲识别提供支持。项目页面:https://github.com/cagries/tipofmyear。
原文摘要 · Abstract (English)
Recognizing jazz standards from audio is a challenging form of tune-level music retrieval: different performances of the same standard may vary in tempo, key, arrangement, instrumentation, improvisational content, and even whether the head melody is present. We study this problem using a curated subset of the Jazz Trio Database designed for cross-performance standard recognition. We compare a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models, using both supervised probing and nearest-neighbor retrieval. Our results suggest that from-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-$k$ results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. Our findings motivate jazz standard recognition as a useful stress test for music representation models and as a step toward retrieval-based standard identification. Project page: https://github.com/cagries/tipofmyear.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。