arXiv:2509.07756cs.SDcs.AI2025-09被引 2

对比多种音频特征,发现梅尔谱图和梅尔倒谱系数最适配深度网络分类。

Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks

  • 用深度CNN端到端训练,比较六种音频特征的分类表现
  • 在ESC-50数据集上,梅尔谱图和MFCC准确率显著领先
  • 适合音频分类任务的特征选择与模型设计参考

除决策树和k近邻算法外,深度卷积神经网络(CNN)广泛用于音乐、语音及环境声音等领域的音频分类。训练特定CNN时,可选用梅尔谱图、梅尔频率倒谱系数(MFCC)、循环暂态图、短时傅里叶变换(STFT)色度图、常数Q变换(CQT)色度图和色度能量归一化统计(CENS)色度图等谱特征与节奏特征作为神经网络的数字图像输入。本研究通过端到端深度学习流程,在包含2000条标注环境音频的ESC-50数据集上,详细评估了这些特征在音频类别级与音频类级分类任务中的性能。多分类评估指标(准确率、精确率、召回率、F1分数)明确显示,梅尔谱图和梅尔频率倒谱系数(MFCC)在使用深度CNN进行音频分类任务时,显著优于其他所研究的谱特征与节奏特征。

原文摘要 · Abstract (English)

Next to decision tree and k-nearest neighbours algorithms deep convolutional neural networks (CNNs) are widely used to classify audio data in many domains like music, speech or environmental sounds. To train a specific CNN various spectral and rhythm features like mel-scaled spectrograms, mel-frequency cepstral coefficients (MFCC), cyclic tempograms, short-time Fourier transform (STFT) chromagrams, constant-Q transform (CQT) chromagrams and chroma energy normalized statistics (CENS) chromagrams can be used as digital image input data for the neural network. The performance of these spectral and rhythm features for audio category level as well as audio class level classification is investigated in detail with a deep CNN and the ESC-50 dataset with 2,000 labeled environmental audio recordings using an end-to-end deep learning pipeline. The evaluated metrics accuracy, precision, recall and F1 score for multiclass classification clearly show that the mel-scaled spectrograms and the mel-frequency cepstral coefficients (MFCC) perform significantly better then the other spectral and rhythm features investigated in this research for audio classification tasks using deep CNNs.

音频分类深度学习特征对比梅尔谱图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。