arXiv:2506.23873cs.SDcs.IR2025-06中稿 · ISMIR 2025被引 6

对比自监督学习让变压器模型自发产生音乐特征,无需专门训练

Emergent musical properties of a transformer under contrastive self-supervised learning

  • 用一维视觉变压器+对比学习,在频时域建模音乐信号
  • 序列标记虽未直接训练,却意外捕捉到和弦等局部音乐信息
  • 不同层注意力图揭示了起音、节奏等高阶音乐特征的涌现

在音乐信息检索中,对比自监督学习虽在全局任务(如自动标签)上有效,但普遍认为其对局部任务(如和弦估计)效果不佳,需更复杂的掩码建模。本文挑战这一观点,采用一维视觉变压器(ViT-1D)在频时域以简单对比学习(NT-Xent)训练。尽管损失函数仅作用于类别标记,但得益于权重共享,序列标记中涌现出丰富音乐属性。全局任务中,类别与序列标记的平均表现优于仅用类别标记;局部任务中,序列标记表现意外出色。此外,逐层注意力图与自相似矩阵显示不同层捕捉不同音乐维度,如起音等高阶特征。本文不追求性能提升,而是揭示变压器在音乐理解中的潜在机制,凸显对比学习与变压器结合在序列建模中的被忽视能力。

原文摘要 · Abstract (English)

In music information retrieval (MIR), contrastive self-supervised learning for general-purpose representation models is effective for global tasks such as automatic tagging. However, for local tasks such as chord estimation, it is widely assumed that contrastively trained general-purpose self-supervised models are inadequate and that more sophisticated SSL is necessary; e.g., masked modeling. Our paper challenges this assumption by revealing the potential of contrastive SSL paired with a transformer in local MIR tasks. We consider a lightweight vision transformer with one-dimensional patches in the time--frequency domain (ViT-1D) and train it with simple contrastive SSL through normalized temperature-scaled cross-entropy loss (NT-Xent). Although NT-Xent operates only over the class token, we observe that, potentially thanks to weight sharing, informative musical properties emerge in ViT-1D's sequence tokens. On global tasks, the temporal average of class and sequence tokens offers a performance increase compared to the class token alone, showing useful properties in the sequence tokens. On local tasks, sequence tokens perform unexpectedly well, despite not being specifically trained for. Furthermore, high-level musical features such as onsets emerge from layer-wise attention maps and self-similarity matrices show different layers capture different musical dimensions. Our paper does not focus on improving performance but advances the musical interpretation of transformers and sheds light on some overlooked abilities of contrastive SSL paired with transformers for sequence modeling in MIR.

音乐生成自监督学习视觉变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。