arXiv:2509.09931eess.AS2025-09被引 4

不依赖知识蒸馏的CNN-GRU模型实现低复杂度音景分类

Acoustic Scene Classification Using CNN-GRU Model Without Knowledge Distillation

  • 用CNN-GRU架构直接训练,无知识蒸馏步骤
  • 内存仅114.2KB,MAC操作10.9M,推理开销极低
  • 在TAU Urban Acoustic Scene 2022数据集上达60.25%准确率

本文介绍SNTL-NTU团队在DCASE 2025低复杂度音景与事件挑战赛任务1的提交方案。该方案摒弃传统教师-学生模型的知识蒸馏范式,旨在以有限计算资源实现高精度。所提模型基于CNN-GRU结构,仅使用TAU Urban Acoustic Scene 2022 Mobile开发数据集进行训练,除用于设备脉冲响应(DIR)增强的MicIRP外未引入外部数据。模型内存占用为114.2KB,乘加运算量(MAC)为10.9M。在开发集上,模型达到60.25%的分类准确率。

原文摘要 · Abstract (English)

In this technical report, we present the SNTL-NTU team's Task 1 submission for the Low-Complexity Acoustic Scenes and Events (DCASE) 2025 challenge. This submission departs from the typical application of knowledge distillation from a teacher to a student model, aiming to achieve high performance with limited complexity. The proposed model is based on a CNN-GRU model and is trained solely using the TAU Urban Acoustic Scene 2022 Mobile development dataset, without utilizing any external datasets, except for MicIRP, which is used for device impulse response (DIR) augmentation. The proposed model has a memory usage of 114.2KB and requires 10.9M muliply-and-accumulate (MAC) operations. Using the development dataset, the proposed model achieved an accuracy of 60.25%.

音景分类CNN-GRU低复杂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。