arXiv:2603.08759cs.SDcs.AI2026-03

针对电音结构分割难题,提出专用于电子舞曲的自监督模型EDMFormer。

EDMFormer: Genre-Specific Self-Supervised Learning for Music Structure Segmentation

  • 基于电音特有的能量、节奏和音色变化设计自监督学习框架
  • 在98首专业标注电音曲目上,对drop和buildup段落检测准确率显著提升
  • 适合音乐信息检索、智能编曲等电音相关应用

音乐结构分割是音频分析的关键任务,但现有模型在电子舞曲(EDM)上的表现不佳。原因在于多数方法依赖歌词或和声相似性,适用于流行音乐却难以捕捉EDM的结构特征。EDM结构由能量、节奏与音色变化定义,包含buildup、drop、breakdown等段落。本文提出EDMFormer,一种结合自监督音频嵌入与专用于电音的数据集及分类体系的Transformer模型。我们发布了EDM-98数据集,包含98首经专业人士标注的电音曲目。实验表明,EDMFormer在边界检测和段落标注方面优于现有模型,尤其在drop和buildup段落识别上表现突出。结果表明,将学习表征与类型特定数据及结构先验相结合,对电音有效,也可推广至其他特殊音乐类型或更广泛的音频领域。

原文摘要 · Abstract (English)

Music structure segmentation is a key task in audio analysis, but existing models perform poorly on Electronic Dance Music (EDM). This problem exists because most approaches rely on lyrical or harmonic similarity, which works well for pop music but not for EDM. EDM structure is instead defined by changes in energy, rhythm, and timbre, with different sections such as buildup, drop, and breakdown. We introduce EDMFormer, a transformer model that combines self-supervised audio embeddings using an EDM-specific dataset and taxonomy. We release this dataset as EDM-98: a group of 98 professionally annotated EDM tracks. EDMFormer improves boundary detection and section labelling compared to existing models, particularly for drops and buildups. The results suggest that combining learned representations with genre-specific data and structural priors is effective for EDM and could be applied to other specialized music genres or broader audio domains.

音乐分割自监督学习电子舞曲Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。