arXiv:2507.13572cs.SDeess.AS2025-07中稿 · WASPAA 2025被引 2

让预训练音乐模型高效分析整首歌,提升结构识别准确率。

Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis

  • 扩展音频窗口并降低时间分辨率,适配长音频分析
  • 在两个数据集上同时提升边界检测与结构功能预测性能
  • 适合需要快速处理完整歌曲的音乐信息检索任务

基于音频的音乐结构分析(MSA)是音乐信息检索中的关键任务,但因音乐形式的复杂性和多样性仍具挑战。近期研究显示,微调预训练音乐基础模型在MSA任务中具有潜力。然而,这些模型通常以高时间分辨率和短音频窗口训练,导致处理长音频时效率低且引入偏差。本文提出一种针对MSA任务的时序适应方法,通过两种关键策略实现全曲单次前向传播的高效分析:(1) 音频窗口扩展,(2) 低分辨率适应。在Harmonix Set和RWC-Pop数据集上的实验表明,该方法显著提升边界检测与结构功能预测性能,同时保持相近的内存占用与推理速度。

原文摘要 · Abstract (English)

Audio-based music structure analysis (MSA) is an essential task in Music Information Retrieval that remains challenging due to the complexity and variability of musical form. Recent advances highlight the potential of fine-tuning pre-trained music foundation models for MSA tasks. However, these models are typically trained with high temporal feature resolution and short audio windows, which limits their efficiency and introduces bias when applied to long-form audio. This paper presents a temporal adaptation approach for fine-tuning music foundation models tailored to MSA. Our method enables efficient analysis of full-length songs in a single forward pass by incorporating two key strategies: (1) audio window extension and (2) low-resolution adaptation. Experiments on the Harmonix Set and RWC-Pop datasets show that our method significantly improves both boundary detection and structural function prediction, while maintaining comparable memory usage and inference speed.

音乐分析时序建模基础模型结构识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。