用多种声学特征堆叠提升CNN在小数据下的声音分类性能。
Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification
- 将多种声学特征(如梅尔谱、频谱对比等)堆叠输入CNN,增强表征能力。
- 在ESC-50和UrbanSound8K上表现优于仅用原始谱图的CNN,接近Transformer模型。
- 适合算力有限或标注数据少的边缘设备部署场景。
环境声音分类(ESC)因在智慧城市监控、故障检测、声学安防和制造质量控制中的广泛应用而受到广泛关注。为提升卷积神经网络(CNN)性能,研究探索了特征堆叠技术,将互补的声学描述符整合为更丰富的输入表示。本文研究了基于CNN的模型,采用多种特征堆叠组合,包括对数梅尔谱(LM)、频谱对比(SPC)、音高(CH)、Tonnetz(TZ)、梅尔频率倒谱系数(MFCCs)和伽马音调倒谱系数(GTCC)。在广泛使用的ESC-50和UrbanSound8K数据集上,通过不同训练策略进行实验,包括在ESC-50上预训练、在UrbanSound8K上微调,以及与在大规模语料库AudioSet上预训练的音频谱图变换器(AST)模型进行比较。该实验设计使我们能够分析在不同训练数据量和预训练多样性条件下,特征堆叠的CNN与基于Transformer的模型之间的性能差异。结果表明,当缺乏大规模预训练或大量训练数据时,特征堆叠的CNN提供了更高效且计算成本更低的替代方案,特别适用于资源受限和边缘级的声音分类任务。
原文摘要 · Abstract (English)
Environmental sound classification (ESC) has gained significant attention due to its diverse applications in smart city monitoring, fault detection, acoustic surveillance, and manufacturing quality control. To enhance CNN performance, feature stacking techniques have been explored to aggregate complementary acoustic descriptors into richer input representations. In this paper, we investigate CNN-based models employing various stacked feature combinations, including Log-Mel Spectrogram (LM), Spectral Contrast (SPC), Chroma (CH), Tonnetz (TZ), Mel-Frequency Cepstral Coefficients (MFCCs), and Gammatone Cepstral Coefficients (GTCC). Experiments are conducted on the widely used ESC-50 and UrbanSound8K datasets under different training regimes, including pretraining on ESC-50, fine-tuning on UrbanSound8K, and comparison with Audio Spectrogram Transformer (AST) models pretrained on large-scale corpora such as AudioSet. This experimental design enables an analysis of how feature-stacked CNNs compare with transformer-based models under varying levels of training data and pretraining diversity. The results indicate that feature-stacked CNNs offer a more computationally and data-efficient alternative when large-scale pretraining or extensive training data are unavailable, making them particularly well suited for resource-constrained and edge-level sound classification scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。