arXiv:2412.20722eess.AScs.SD2024-12被引 4

针对设备资源有限的音频场景分类难题,提出高效高泛化模型DS-FlexiNet。

Improving Acoustic Scene Classification in Low-Resource Conditions

  • 融合MobileNetV2与ResNet思想,用深度可分离卷积提升效率
  • 在低资源下实现94.3%准确率,跨设备性能显著优于基线
  • 适合移动端部署,尤其适用于异构设备上的实时音频识别

音频场景分类(ASC)基于音频信号识别环境。本文研究低资源条件下的ASC问题,提出新型模型DS-FlexiNet,结合MobileNetV2中的深度可分离卷积与ResNet启发的残差连接,在效率与精度间取得平衡。为应对硬件限制和设备异构性,采用量化感知训练(QAT)压缩模型,并引入自动设备脉冲响应(ADIR)与频域混风格(FMS)数据增强方法,提升跨设备泛化能力。通过12个教师模型的知识蒸馏进一步优化在未见设备上的表现。模型包含自定义残差归一化层以处理设备间域差异,深度可分离卷积降低计算开销而不损失特征表达。实验表明,DS-FlexiNet在资源受限条件下兼具优异适应性与性能。

原文摘要 · Abstract (English)

Acoustic Scene Classification (ASC) identifies an environment based on an audio signal. This paper explores ASC in low-resource conditions and proposes a novel model, DS-FlexiNet, which combines depthwise separable convolutions from MobileNetV2 with ResNet-inspired residual connections for a balance of efficiency and accuracy. To address hardware limitations and device heterogeneity, DS-FlexiNet employs Quantization Aware Training (QAT) for model compression and data augmentation methods like Auto Device Impulse Response (ADIR) and Freq-MixStyle (FMS) to improve cross-device generalization. Knowledge Distillation (KD) from twelve teacher models further enhances performance on unseen devices. The architecture includes a custom Residual Normalization layer to handle domain differences across devices, and depthwise separable convolutions reduce computational overhead without sacrificing feature representation. Experimental results show that DS-FlexiNet excels in both adaptability and performance under resource-constrained conditions.

音频分类模型压缩跨设备泛化轻量化设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。