arXiv:2501.01921cs.SDeess.AS2025-01

融合音效结构与统计特征,提升环境声音分类精度

Structural and Statistical Audio Texture Knowledge Distillation for Acoustic Classification

  • 从中间层提取低频音效纹理,结合高层语义信息
  • 在四个数据集上均实现稳定准确率提升
  • 适合需要高精度声音识别的智能监控场景

尽管知识蒸馏在多种音频任务中表现良好,但在环境声音分类中常忽略捕捉复杂声学环境中局部模式所必需的低级音频纹理特征。为此,本文提出结构与统计音频纹理知识蒸馏框架(SSATKD),将高层上下文信息与从中间层提取的低级结构和统计音频纹理相结合。为验证其在不同声学领域的泛化能力,SSATKD在四个环境声音分类数据集上进行测试,包括两个被动声纳数据集(DeepShip 和 VTUAD)以及两个通用环境声音数据集(ESC-50 和 TUT 音景数据集)。探索了两种教师模型适应策略:仅适配分类头与全微调。同时使用多种卷积和基于变压器的教师模型进行评估。实验结果表明,在所有数据集和设置下均实现一致的准确率提升,证实了 SSATKD 在真实世界声音分类任务中的有效性和鲁棒性。

原文摘要 · Abstract (English)

While knowledge distillation has shown success in various audio tasks, its application to environmental sound classification often overlooks essential low-level audio texture features needed to capture local patterns in complex acoustic environments. To address this gap, the Structural and Statistical Audio Texture Knowledge Distillation (SSATKD) framework is proposed, which combines high-level contextual information with low-level structural and statistical audio textures extracted from intermediate layers. To evaluate its generalizability across diverse acoustic domains, SSATKD is tested on four datasets within the environmental sound classification domain, including two passive sonar datasets (DeepShip and Vessel Type Underwater Acoustic Data (VTUAD)) and two general environmental sound datasets (Environmental Sound Classification 50 (ESC-50) and Tampere University of Technology (TUT) Acoustic Scenes). Two teacher adaptation strategies are explored: classifier-head-only adaptation and full fine-tuning. The framework is further evaluated using various convolutional and transformer-based teacher models. Experimental results demonstrate consistent accuracy improvements across all datasets and settings, confirming the effectiveness and robustness of SSATKD in real-world sound classification tasks.

音频分类知识蒸馏声纹分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。