arXiv:2512.17209cs.SDcs.LG2025-12中稿 · ICASSP 2026被引 2

探究预训练音频模型如何理解音乐结构

Do Foundational Audio Encoders Understand Music Structure?

  • 对比11种预训练音频模型,分析其音乐结构解析能力
  • 自监督学习+掩码语言建模的模型在结构分析中表现最佳
  • 为音频模型与音乐分析融合提供新方向

在音乐信息检索领域,预训练基础音频编码器(FAEs)已成为研究趋势。这些在大量音乐和音频数据上预训练的模型,在音乐标签和自动音乐转录等任务中表现出色。然而,其在音乐结构分析(MSA)中的应用仍不充分:仅有少量FAE被用于MSA,且学习方法、训练数据、模型上下文长度等因素对性能的影响尚不明确。本研究对11类FAEs进行了全面实验,探讨这些因素如何影响MSA表现。结果表明,基于音乐数据进行自监督学习并采用掩码语言建模的FAEs在音乐结构分析中尤为有效。这一发现为未来FAE与音乐结构分析的研究提供了新路径。

原文摘要 · Abstract (English)

In music information retrieval (MIR) research, the use of pretrained foundational audio encoders (FAEs) has recently become a trend. FAEs pretrained on large amounts of music and audio data have been shown to improve performance on MIR tasks such as music tagging and automatic music transcription. However, their use for music structure analysis (MSA) remains underexplored: only a small subset of FAEs has been examined for MSA, and the impact of factors such as learning methods, training data, and model context length on MSA performance remains unclear. In this study, we conduct comprehensive experiments on 11 types of FAEs to investigate how these factors affect MSA performance. Our results demonstrate that FAEs using self-supervised learning with masked language modeling on music data are particularly effective for MSA. These findings pave the way for future research in FAE and MSA.

音乐结构音频编码器自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。