arXiv:2603.27218cs.SDcs.AI2026-03

用无监督方法评估9个音频模型在音乐结构分析中的表现

Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis

  • 用三种无监督算法对音符级嵌入进行分割,仅关注边界检测
  • 现代通用音频嵌入优于传统谱图基方法,但效果不一致
  • CBM算法最有效,建议采用剪裁标注提升评估严谨性

音乐结构分析(MSA)旨在揭示乐曲的高层组织结构。现有先进方法多依赖有监督深度学习,但受限于大量标注数据及结构本身的模糊性。本文提出一种无监督评估方法,对九种开源通用预训练音频模型在MSA上的表现进行评估。针对每个模型,提取每小节的嵌入表示,并使用三种无监督分割算法(Foote的棋盘核、谱聚类、相关块匹配法(CBM))进行分割,专注于边界检测任务。结果表明,现代通用深度嵌入整体优于传统基于谱图的基线方法,但并非系统性更优。此外,本研究提出的无监督边界估计方法性能普遍优于近期线性探针基线。在所评估的方法中,CBM算法始终表现最佳。最后,我们指出标准评估指标存在人为膨胀问题,倡导系统采用‘剪裁’甚至‘双重剪裁’标注以建立更严格的MSA评估标准。

原文摘要 · Abstract (English)

Music Structure Analysis (MSA) aims to uncover the high-level organization of musical pieces. State-of-the-art methods are often based on supervised deep learning, but these methods are bottlenecked by the need for heavily annotated data and inherent structural ambiguities. In this paper, we propose an unsupervised evaluation of nine open-source, generic pre-trained deep audio models, on MSA. For each model, we extract barwise embeddings and segment them using three unsupervised segmentation algorithms (Foote's checkerboard kernels, spectral clustering, and Correlation Block-Matching (CBM)), focusing exclusively on boundary retrieval. Our results demonstrate that modern, generic deep embeddings generally outperform traditional spectrogram-based baselines, but not systematically. Furthermore, our unsupervised boundary estimation methodology generally yields stronger performance than recent linear probing baselines. Among the evaluated techniques, the CBM algorithm consistently emerges as the most effective downstream segmentation method. Finally, we highlight the artificial inflation of standard evaluation metrics and advocate for the systematic adoption of ``trimming'', or even ``double trimming'' annotations to establish more rigorous MSA evaluation standards.

音乐分析无监督学习音频嵌入结构检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。