用深度卷积网络对比音频特征,发现梅尔谱图和梅尔倒谱系数效果最好。
Spectral and Rhythm Features for Audio Classification with Deep Convolutional Neural Networks

- 采用深度CNN比较多种音频特征表示方法。
- 梅尔谱图和梅尔倒谱系数在分类任务中表现最优,准确率显著更高。
- 适合音频分类研究者参考,尤其关注特征工程的实践者。
卷积神经网络(CNN)广泛应用于计算机视觉领域,不仅可用于图像模式识别,也可用于从时域数字音频信号中提取频谱与节奏特征,实现声音分类。本文研究了多种频谱与节奏特征表示方法在深度卷积神经网络上的音频分类性能,包括梅尔尺度谱图、梅尔频率倒谱系数(MFCCs)、循环临时图、短时傅里叶变换(STFT)色度图、恒Q变换(CQT)色度图以及色度能量归一化统计(CENS)色度图。实验表明,梅尔谱图和梅尔倒谱系数在使用深度CNN进行音频分类任务中表现显著优于其他特征。实验基于包含2000个标注环境音频记录的ESC-50数据集开展。
原文摘要 · Abstract (English)
Convolutional neural networks (CNNs) are widely used in computer vision. They can be used not only for conventional digital image material to recognize patterns, but also for feature extraction from digital imagery representing spectral and rhythm features extracted from time-domain digital audio signals for the acoustic classification of sounds. Different spectral and rhythm feature representations like mel-scaled spectrograms, mel-frequency cepstral coefficients (MFCCs), cyclic tempograms, short-time Fourier transform (STFT) chromagrams, constant-Q transform (CQT) chromagrams and chroma energy normalized statistics (CENS) chromagrams are investigated in terms of the audio classification performance using a deep convolutional neural network. It can be clearly shown that the mel-scaled spectrograms and the mel-frequency cepstral coefficients (MFCCs) perform significantly better than the other spectral and rhythm features investigated in this research for audio classification tasks using deep CNNs. The experiments were carried out with the aid of the ESC-50 dataset with 2,000 labeled environmental audio recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。