arXiv:2511.09802eess.SPcs.SD2025-11

提出稀疏池化方法,显著提升环境音分类精度与效率

Investigation of Feature Selection and Pooling Methods for Environmental Sound Classification

  • 采用稀疏显著区域池化及其变体,降低特征维度
  • SSRP-T在ESC-50上达80.69%准确率,远超基线66.75%和PCA模型37.60%
  • 适合边缘设备部署,兼顾精度与计算开销

本文研究轻量级CNN中降维与池化方法对环境音分类(ESC)的影响。评估了稀疏显著区域池化(SSRP)及其变体SSRP-Basic(SSRP-B)和SSRP-Top-K(SSRP-T)在不同超参数设置下的表现,并与主成分分析(PCA)对比。在ESC-50数据集上的实验表明,SSRP-T最高可达80.69%准确率,显著优于基准CNN(66.75%)和PCA降维模型(37.60%)。结果证实,经过调优的稀疏池化策略为ESC任务提供了鲁棒、高效且高性能的解决方案,尤其适用于需平衡精度与计算成本的资源受限场景。

原文摘要 · Abstract (English)

This paper explores the impact of dimensionality reduction and pooling methods for Environmental Sound Classification (ESC) using lightweight CNNs. We evaluate Sparse Salient Region Pooling (SSRP) and its variants, SSRP-Basic (SSRP-B) and SSRP-Top-K (SSRP-T), under various hyperparameter settings and compare them with Principal Component Analysis (PCA). Experiments on the ESC-50 dataset demonstrate that SSRP-T achieves up to 80.69 % accuracy, significantly outperforming both the baseline CNN (66.75 %) and the PCA-reduced model (37.60 %). Our findings confirm that a well-tuned sparse pooling strategy provides a robust, efficient, and high-performing solution for ESC tasks, particularly in resource-constrained scenarios where balancing accuracy and computational cost is crucial.

环境音分类稀疏池化轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。