arXiv:2604.05011cs.SDcs.AI2026-04被引 1

构建首个也门传统音乐分类数据集与模型,准确率达98.8%。

YMIR: A new Benchmark Dataset and Model for Arabic Yemeni Music Genre Classification Using Convolutional Neural Networks

  • 基于卷积神经网络设计专用分类模型,利用时频特征进行识别
  • 在1475段音频上实现98.8%最高准确率,梅尔谱图表现最优
  • 为阿拉伯也门音乐研究提供可复现基准,适合文化音源分析者

自动音乐流派分类是音乐信息检索中的关键任务,但现有大多数基准和模型主要针对西方音乐,忽视了文化特异性传统。本文提出也门音乐信息检索(YMIR)数据集,包含1475段精心挑选的音频片段,覆盖萨那、哈德拉米、拉吉、蒂哈米和阿德尼五种传统也门流派。数据由五名也门音乐专家按标准化协议标注,达成高一致性(Fleiss kappa = 0.85)。同时提出也门音乐分类模型(YMCM),一种基于卷积神经网络(CNN)的系统,用于从时频特征中分类音乐流派。采用一致预处理流程,在六个实验组和五种架构下共完成30次实验,评估了梅尔谱图、染色体、滤波器组及梅尔频率倒谱系数(MFCCs,13/20/40系数)等多种特征表示,并在相同条件下对比了标准模型(AlexNet、VGG16、MobileNet)与基线CNN。实验表明,YMCM表现最佳,使用梅尔谱图特征时准确率达98.8%。结果揭示了特征表示与模型容量之间的实际关系。研究确立了YMIR作为有用基准,以及YMCM作为也门音乐流派分类的强基线。

原文摘要 · Abstract (English)

Automatic music genre classification is a major task in music information retrieval; however, most current benchmarks and models have been developed primarily for Western music, leaving culturally specific traditions underrepresented. In this paper, we introduce the Yemeni Music Information Retrieval (YMIR) dataset, which contains 1,475 carefully selected audio clips covering five traditional Yemeni genres: Sanaani, Hadhrami, Lahji, Tihami, and Adeni. The dataset was labeled by five Yemeni music experts following a clear and structured protocol, resulting in strong inter-annotator agreement (Fleiss kappa = 0.85). We also propose the Yemeni Music Classification Model (YMCM), a convolutional neural network (CNN)-based system designed to classify music genres from time-frequency features. Using a consistent preprocessing pipeline, we perform a systematic comparison across six experimental groups and five different architectures, resulting in a total of 30 experiments. Specifically, we evaluate several feature representations, including Mel-spectrograms, Chroma, FilterBank, and MFCCs with 13, 20, and 40 coefficients, and benchmark YMCM against standard models (AlexNet, VGG16, MobileNet, and a baseline CNN) under the same experimental conditions. The experimental findings reveal that YMCM is the most effective, achieving the highest accuracy of 98.8% with Mel-spectrogram features. The results also provide practical insights into the relationship between feature representation and model capacity. The findings establish YMIR as a useful benchmark and YMCM as a strong baseline for classifying Yemeni music genres.

音乐分类数据集深度学习文化音源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。