SongFormer通过异构监督实现音乐结构分析的规模化,提升精度与鲁棒性。
SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
- 融合长短窗自监督表示,捕捉细粒度与长程依赖
- 在300首专家验证曲目上达到最高功能标签准确率
- 适合音乐理解、生成及跨语言跨风格研究者
音乐结构分析(MSA)是音乐理解与可控生成的基础,但受限于小规模且不一致的数据集。本文提出SongFormer,一个可扩展的框架,能从异构监督中学习。SongFormer (i) 融合短窗与长窗自监督学习表征,以捕捉细粒度与长程依赖;(ii) 引入可学习的源嵌入,支持部分、噪声大及模式不匹配标签的训练。为支持规模化与公平评估,我们发布了目前最大的MSA数据集SongFormDB(超过14,000首歌曲,涵盖多种语言与风格),以及300首专家验证的SongFormBench基准。在SongFormBench上,SongFormer在严格边界检测(HR.5F)上达到新最优,并取得最高功能标签准确率,同时保持计算高效;其性能超越多个强基线模型及Gemini 2.5 Pro,在宽松容忍度下(HR3F)仍具竞争力。代码、数据集与模型已开源。
原文摘要 · Abstract (English)
Music structure analysis (MSA) underpins music understanding and controllable generation, yet progress has been limited by small, inconsistent corpora. We present SongFormer, a scalable framework that learns from heterogeneous supervision. SongFormer (i) fuses short- and long-window self-supervised learning representations to capture both fine-grained and long-range dependencies, and (ii) introduces a learned source embedding to enable training with partial, noisy, and schema-mismatched labels. To support scaling and fair evaluation, we release SongFormDB, the largest MSA corpus to date (over 14k songs spanning languages and genres), and SongFormBench, a 300-song expert-verified benchmark. On SongFormBench, SongFormer sets a new state of the art in strict boundary detection (HR.5F) and achieves the highest functional label accuracy, while remaining computationally efficient; it surpasses strong baselines and Gemini 2.5 Pro on these metrics and remains competitive under relaxed tolerance (HR3F). Code, datasets, and model are open-sourced at https://github.com/ASLP-lab/SongFormer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。