arXiv:2604.10021cs.SDcs.LG2026-04

用掩码对比预训练提升音乐调性检测,效果超越现有方法。

Masked Contrastive Pre-Training Improves Music Audio Key Detection

论文配图:Masked Contrastive Pre-Training Improves Music Audio Key Detection
图 1 · 摘自论文原文
  • 通过掩码对比学习提取音高敏感特征,无需复杂数据增强。
  • 在音乐调性检测任务上达到当前最佳(SOTA)性能。
  • 适合关注自监督音乐模型设计的研究者与开发者。

自监督音乐基础模型在调性检测任务上表现不佳,因其需要对音高敏感的表征。本文首次系统性研究发现,自监督预训练的设计直接影响音高敏感性,并证明掩码对比嵌入在监督设置下能实现最先进的(SOTA)调性检测性能。首先,我们发现基于梅尔频谱图的掩码对比预训练后进行线性评估,即可直接获得具有竞争力的调性检测表现。这促使我们使用浅而宽的多层感知机(MLPs)在基础模型提取的特征上进行训练,从而在无需复杂数据增强策略的情况下实现SOTA性能。进一步分析表明,所学表征天然编码了常见数据增强方式。本研究确立了自监督预训练在音高敏感音乐信息检索任务中的有效性,并为音乐基础模型的设计与探测提供了洞见。

原文摘要 · Abstract (English)

Self-supervised music foundation models underperform on key detection, which requires pitch-sensitive representations. In this work, we present the first systematic study showing that the design of self-supervised pretraining directly impacts pitch sensitivity, and demonstrate that masked contrastive embeddings uniquely enable state-of-the-art (SOTA) performance in key detection in the supervised setting. First, we discover that linear evaluation after masking-based contrastive pretraining on Mel spectrograms leads to competitive performance on music key detection out of the box. This leads us to train shallow but wide multi-layer perceptrons (MLPs) on features extracted from our base model, leading to SOTA performance without the need for sophisticated data augmentation policies. We further analyze robustness and show empirically that the learned representations naturally encode common augmentations. Our study establishes self-supervised pretraining as an effective approach for pitch-sensitive MIR tasks and provides insights for designing and probing music foundation models.

音乐信息检索自监督学习调性检测对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。