arXiv:2506.17055cs.SDcs.IR2025-06中稿 · ISMIR 2025被引 6

评测五大音频基础模型在六大世界音乐数据集上的表现,揭示其跨文化泛化能力。

Universal Music Representations? Evaluating Foundation Models on World Music Corpora

  • 用探针、微调和少样本学习三种方法评估模型跨文化表现。
  • 五组数据达当前最优,大模型在非西方音乐上表现更优但文化差异影响效果。
  • 微调不总优于探针,说明模型已具备丰富音乐知识,适合研究全球音乐理解者。

音频基础模型革新了音乐信息检索,但其在多元音乐传统间的泛化能力仍存疑问。本文对五种顶尖音频基础模型在涵盖西方流行乐、希腊、土耳其及印度古典音乐的六个音乐语料库上进行了全面评估。采用三种互补方法:探针分析内在表征、1-2层针对性监督微调、多标签少样本学习以应对低资源场景。分析显示,不同文化间泛化能力存在差异,大模型在非西方音乐中通常表现更佳,但在文化差异较大的传统中性能下降。值得注意的是,所提方法在六项评估数据集中有五项达到当前最优水平,证明基础模型在世界音乐理解中的有效性。此外,目标微调并未在所有情况下超越探针,表明基础模型已编码大量音乐知识。本研究的评估框架与基准结果有助于理解当前模型距离通用音乐表征还有多远,并为未来进展提供衡量标准。

原文摘要 · Abstract (English)

Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of-the-art audio foundation models across six musical corpora spanning Western popular, Greek, Turkish, and Indian classical traditions. We employ three complementary methodologies to investigate these models' cross-cultural capabilities: probing to assess inherent representations, targeted supervised fine-tuning of 1-2 layers, and multi-label few-shot learning for low-resource scenarios. Our analysis shows varying cross-cultural generalization, with larger models typically outperforming on non-Western music, though results decline for culturally distant traditions. Notably, our approaches achieve state-of-the-art performance on five out of six evaluated datasets, demonstrating the effectiveness of foundation models for world music understanding. We also find that our targeted fine-tuning approach does not consistently outperform probing across all settings, suggesting foundation models already encode substantial musical knowledge. Our evaluation framework and benchmarking results contribute to understanding how far current models are from achieving universal music representations while establishing metrics for future progress.

音乐理解基础模型跨文化少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。