arXiv:2509.11474cs.SDcs.IR2025-09

发现电子舞曲商业分类远超真实声学差异,存在近半数冗余标签。

Acoustic Overspecification in Electronic Dance Music Taxonomy

  • 构建专用于电子舞曲的可解释声学特征空间,捕捉制作技法与节奏纹理。
  • 无论自定义特征或预训练音频嵌入,聚类均发现不超过20个自然声学类别。
  • 揭示当前商业分类在声学上过度细分,适合音乐信息检索与音乐认知研究者。

电子舞曲(EDM)分类通常依赖产业定义的分类体系,现有监督方法默认这些子流派标签有效。然而,这些商业区分是否反映真实的声学差异仍不清楚。本文提出一种无监督方法,以发现独立于商业标签的自然声学结构。为弥补音乐信息检索领域缺乏针对EDM的特征设计,我们系统构建了一个定制的、可解释的声学特征空间,涵盖该类型的标志性制作技术、频谱质感及层叠节奏模式。为确保结果反映内在声学结构而非特征工程偏差,我们在状态前沿的预训练音频嵌入(MERT与CLAP)上验证聚类效果。无论在自定义特征空间还是预训练嵌入中,聚类始终识别出20个或更少的自然声学家族,表明当前商业EDM分类在声学上被过度指定,接近一半的标签并不对应真实的声学差异。

原文摘要 · Abstract (English)

Electronic Dance Music (EDM) classification typically relies on industry-defined taxonomies, with current supervised approaches naturally assuming the validity of prescribed subgenre labels. However, whether these commercial distinctions reflect genuine acoustic differences remains largely unexplored. In this paper, we propose an unsupervised approach to discover the natural acoustic structure of EDM independent of commercial labels. To address the historical lack of EDM-specific feature design in MIR, we systematically construct a tailored, interpretable acoustic feature space capturing the genre's defining production techniques, spectral textures, and layered rhythmic patterns. To ensure our findings reflect inherent acoustic structure rather than feature engineering artifacts, we validate our clustering against state-of-the-art pre-trained audio embeddings (MERT and CLAP). Across both our bespoke feature space and the pre-trained embeddings, clustering consistently identifies 20 or fewer natural acoustic families -- suggesting current commercial EDM taxonomy is acoustically overspecified by nearly one-half.

电子舞曲声学分类无监督学习音乐信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。