arXiv:2501.12907cs.SDeess.AS2025-01被引 9

无需人工标注,通过自监督学习区分音乐的大调与小调。

S-KEY: Self-supervised Learning of Major and Minor Keys from Audio

  • 利用不变于转调的音高特征生成伪标签,设计辅助预训练任务。
  • 在FMAKv2和GTZAN数据集上达到有监督方法的顶尖水平。
  • 可扩展至百万首歌曲,适合大规模音乐信息检索研究者。

当前自监督音乐调性估计方法STONE无法区分相对调(如C大调与A小调)。本文提出S-KEY,通过扩展STONE的神经网络架构与学习目标,实现大调与小调的自监督学习。核心贡献是引入一个基于转调不变音高特征的辅助预训练任务,作为伪标签来源。S-KEY在FMAKv2和GTZAN数据集上的调性估计性能达到有监督方法的最先进水平,且无需人工标注,参数量与STONE相同。在此基础上,我们还将S-KEY训练集扩展至百万首歌曲,展示了大规模自监督学习在音乐信息检索中的潜力。

原文摘要 · Abstract (English)

STONE, the current method in self-supervised learning for tonality estimation in music signals, cannot distinguish relative keys, such as C major versus A minor. In this article, we extend the neural network architecture and learning objective of STONE to perform self-supervised learning of major and minor keys (S-KEY). Our main contribution is an auxiliary pretext task to STONE, formulated using transposition-invariant chroma features as a source of pseudo-labels. S-KEY matches the supervised state of the art in tonality estimation on FMAKv2 and GTZAN datasets while requiring no human annotation and having the same parameter budget as STONE. We build upon this result and expand the training set of S-KEY to a million songs, thus showing the potential of large-scale self-supervised learning in music information retrieval.

自监督学习音乐分析调性识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。