提出新指标选音乐模型最佳层,小数据下比传统方法更准
What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models

- 用内在几何性质分析12个音乐模型各层表示质量
- 发现标准指标对调性任务无效,但新等变性指标有效
- 无需微调,直接用内在指标选层,适合数据少的场景
音乐基础模型常作为冻结的音频特征提取器,但选择哪一层提取仍依赖经验。当前做法多采用固定深度或多层融合,对为何某些层在下游任务中表现更好、表征质量如何随深度和预训练范式变化缺乏理解。我们系统分析了12个音乐基础模型(涵盖掩码建模、自回归建模和对比学习三种预训练范式),通过内在几何与变换特性刻画其隐藏表示。将无标签表征质量度量与15个下游任务的层级性能相关联,发现多个度量能追踪风格分类、情绪识别、自动标签和节拍追踪的任务表现,但对调性任务(如调性估计、和弦识别)均失效,表明无单一属性可通用代表音乐信息检索任务的表征质量。为此,我们引入一种音高平移等变性度量,捕捉现有标准指标未覆盖的特性,为跨模型族提供一致的调性质量指示。最终结果表明,内在度量可作为有效的层选择代理,匹配或超越可训练多层融合方法,尤其在数据有限场景下表现更优。
原文摘要 · Abstract (English)
Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi-layer fusion, with limited understanding of why certain layers transfer better across downstream tasks or how representation quality varies with depth and pre-training paradigm. We conduct a systematic layer-wise analysis of 12 music foundation models spanning three pre-training paradigms (masked modeling, autoregressive modeling, and contrastive learning), characterizing their hidden representations through intrinsic geometric and transformation-based properties. Correlating label-free representation-quality metrics with layer-wise performance across 15 downstream tasks, we find that several metrics track layer quality for genre classification, emotion recognition, automatic tagging, and beat tracking, albeit with varying strength across tasks and pre-training paradigms. However, all metrics fail on tonal tasks such as key estimation and chord recognition, indicating that no single property serves as a general proxy for representation quality across music information retrieval tasks. To address this gap, we introduce a pitch-transposition equivariance measure that captures properties missed by these standard metrics, providing a consistent indicator of tonal quality across model families. Finally, we show that intrinsic metrics can serve as effective proxies for layer selection, matching or outperforming trainable multi-layer fusion methods, particularly in limited-data settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。