arXiv:2506.02091cs.SDcs.LG2025-06

对比不同频谱图缩放方法在多标签音乐流派识别中的效果

Comparison of spectrogram scaling in multi-label Music Genre Recognition

  • 测试多种频谱图缩放策略提升多标签流派识别性能
  • 在超1.8万条手动标注数据上验证方法有效性
  • 适合关注音频特征预处理的音乐信息检索研究者

随着数字音频工作站的普及,普通听众可获取的音乐数量大幅增加;同时,不同音乐流派之间的界限往往模糊且抽象,单张专辑中常出现多种流派的混合。本文描述并比较了多种预处理方法与模型训练策略,以应对当今专辑的多元风格特征。实验基于一个包含超过18000个条目的自建手动标注数据集进行。

原文摘要 · Abstract (English)

As the accessibility and ease-of-use of digital audio workstations increases, so does the quantity of music available to the average listener; additionally, differences between genres are not always well defined and can be abstract, with widely varying combinations of genres across individual records. In this article, multiple preprocessing methods and approaches to model training are described and compared, accounting for the eclectic nature of today's albums. A custom, manually labeled dataset of more than 18000 entries has been used to perform the experiments.

音乐流派识别频谱图多标签分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。