arXiv:2509.02471cs.SDcs.LG2025-09中稿 · IEEE Signal Proces…被引 2

用双分支Mamba模型提升工业异常声音检测精度

ESTM: An Enhanced Dual-Branch Spectral-Temporal Mamba for Anomalous Sound Detection

  • 双路径架构分离时频特征,用SSM捕捉长程依赖
  • 在DCASE 2020数据集上显著提升异常检测准确率
  • 适合需要高灵敏度声音监测的工业场景

工业设备异常声音检测的核心挑战在于建模声学特征的时间-频率耦合特性。现有方法受限于局部感受野,难以捕捉长时间序列模式和跨频带动态耦合效应。本文提出新型框架ESTM,基于双路径Mamba架构,实现时频解耦建模,并采用选择性状态空间模型(SSM)进行长序列建模。ESTM通过融合增强型梅尔谱图与原始音频特征,从不同时间片段和频率带中提取丰富特征表示,同时利用三统计门控模块(TSG)进一步提升对异常模式的敏感性。实验表明,ESTM在DCASE 2020 Task 2数据集上显著提升了异常检测性能,验证了该方法的有效性。

原文摘要 · Abstract (English)

The core challenge in industrial equipment anoma lous sound detection (ASD) lies in modeling the time-frequency coupling characteristics of acoustic features. Existing modeling methods are limited by local receptive fields, making it difficult to capture long-range temporal patterns and cross-band dynamic coupling effects in machine acoustic features. In this paper, we propose a novel framework, ESTM, which is based on a dual-path Mamba architecture with time-frequency decoupled modeling and utilizes Selective State-Space Models (SSM) for long-range sequence modeling. ESTM extracts rich feature representations from different time segments and frequency bands by fusing enhanced Mel spectrograms and raw audio features, while further improving sensitivity to anomalous patterns through the TriStat-Gating (TSG) module. Our experiments demonstrate that ESTM improves anomalous detection performance on the DCASE 2020 Task 2 dataset, further validating the effectiveness of the proposed method.

异常检测音频分析Mamba时频建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。