arXiv:2603.08154cs.SDcs.MM2026-03

用频谱图+CNN提升南亚环境声多标签分类准确率

Soundscapes in Spectrograms: Pioneering Multilabel Classification for South Asian Sounds

  • 基于频谱图设计CNN模型,捕捉复杂重叠声纹
  • 在SAS-KIIT和UrbanSound8K上均显著优于传统MFCC方法
  • 适合城市监测与文化声景分析场景

环境声分类在城市监控与文化声景分析中日益重要,尤其在南亚这类声学环境丰富的地区。该区域常出现自然、人为及文化声音重叠,传统依赖梅尔频率倒谱系数(MFCC)的方法难以应对。本文提出一种新型频谱图基方法,采用卷积神经网络(CNN)解决SAS-KIIT数据集上的多标签多分类难题,并在著名UrbanSound8K数据集上验证其鲁棒性与可比性。结果表明,该方法在两个数据集上均显著优于现有MFCC技术,分类准确率明显提升,为真实应用场景下的音频分类系统提供了更可靠基础。

原文摘要 · Abstract (English)

Environmental sound classification is a field of growing importance for urban monitoring and cultural soundscape analysis, especially within the acoustically rich environments of South Asia. These regions present a unique challenge as multiple natural, human, and cultural sounds often overlap, straining traditional methods that frequently rely on Mel Frequency Cepstral Coefficients (MFCC). This study introduces a novel spectrogram-based methodology with a superior ability to capture these complex auditory patterns. A Convolutional Neural Network (CNN) architecture is implemented to solve a demanding multilabel, multiclass classification problem on the SAS-KIIT dataset. To demonstrate robustness and comparability, the approach is also validated using the renowned UrbanSound8K dataset. The results confirm that the proposed spectrogram-based method significantly outperforms existing MFCC-based techniques, achieving higher classification accuracy across both datasets. This improvement lays the groundwork for more robust and accurate audio classification systems in real-world applications.

声学分类频谱图多标签CNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。