提出自适应时序池化方法,提升无训练异常声音检测效果
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
- 设计相对偏差池化,动态加权异常程度高的音段
- 在5个数据集上超越均值池化,达当前最优性能
- 适合无需微调的工业异常声音实时监测场景
基于预训练音频嵌入的无训练异常声音检测(ASD)近年来受到广泛关注,因其仅需正常参考数据即可检测异常声音,无需任务特定模型训练或微调。然而,现有嵌入式方法几乎全部依赖时序均值池化,时序池化在无训练ASD中尚未被充分探索。本文首次系统评估了预训练音频嵌入在无训练ASD中的时序池化策略。提出相对偏差池化(RDP),一种自适应池化方法,对具有更强时序偏离的嵌入赋予更大权重;研究使用广义均值(GeM)池化的特征级非线性聚合;并考察两者的混合组合。在五个基准数据集上的实验表明,所提池化策略持续优于均值池化,并在无训练ASD中达到最先进水平,包括在DCASE2025 ASD数据集上超越此前报告的训练系统和集成模型。
原文摘要 · Abstract (English)
Training-free anomalous sound detection (ASD) based on pre-trained audio embedding models has recently garnered significant attention, as it enables the detection of anomalous sounds using only normal reference data without task-specific model training or fine-tuning. However, existing embedding-based approaches almost exclusively rely on temporal mean pooling, leaving temporal pooling in training-free ASD largely unexplored. In this paper, we present the first systematic evaluation of temporal pooling strategies for training-free ASD with pre-trained audio embeddings. We propose relative deviation pooling (RDP), an adaptive pooling method that assigns larger weights to embeddings with stronger temporal deviations, investigate feature-wise non-linear aggregation using generalized mean (GeM) pooling, and examine a hybrid combination of both strategies. Experiments on five benchmark datasets demonstrate that the proposed pooling strategies consistently outperform mean pooling and achieve state-of-the-art performance for training-free ASD, including results that surpass previously reported trained systems and ensembles on the DCASE2025 ASD dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。