arXiv:2505.18980cs.SDeess.AS2025-05被引 1

无标签时通过伪异常数据筛选与标注,提升声音异常检测准确率

Improving Anomalous Sound Detection through Pseudo-anomalous Set Selection and Pseudo-label Utilization under Unlabeled Conditions

  • 从外部音频中筛选相似正常声,构建伪异常集
  • 用三元组学习为无标签数据生成伪标签,识别细微异常
  • 迭代优化选择与标注,适合工业场景少样本部署

本文针对在缺乏足够同类设备数据和运行状态标签时,异常声音检测(ASD)性能下降的问题,提出一个整合三阶段流程的端到端方法。首先,采用基于异常得分的选择器,从外部音频中筛选出与目标设备正常声相近的数据;其次,利用三元组学习为未标注数据分配伪标签,实现对运行声音的细粒度分类与微小异常检测;最后,通过迭代训练逐步优化伪异常集选择与伪标签分配,持续提升检测精度。在DCASE2022-2024 Task 2数据集上的实验表明,在无标签条件下,本方法平均AUC提升超过6.6个百分点;在有标签条件下,引入伪异常集中的外部数据进一步增强性能。结果验证了该方法在设备数据稀缺、标注成本高的工业场景下的实用性与鲁棒性,可显著降低人工标注需求。

原文摘要 · Abstract (English)

This paper addresses performance degradation in anomalous sound detection (ASD) when neither sufficiently similar machine data nor operational state labels are available. We present an integrated pipeline that combines three complementary components derived from prior work and extends them to the unlabeled ASD setting. First, we adapt an anomaly score based selector to curate external audio data resembling the normal sounds of the target machine. Second, we utilize triplet learning to assign pseudo-labels to unlabeled data, enabling finer classification of operational sounds and detection of subtle anomalies. Third, we employ iterative training to refine both the pseudo-anomalous set selection and pseudo-label assignment, progressively improving detection accuracy. Experiments on the DCASE2022-2024 Task 2 datasets demonstrate that, in unlabeled settings, our approach achieves an average AUC increase of over 6.6 points compared to conventional methods. In labeled settings, incorporating external data from the pseudo-anomalous set further boosts performance. These results highlight the practicality and robustness of our methods in scenarios with scarce machine data and labels, facilitating ASD deployment across diverse industrial settings with minimal annotation effort.

异常检测声音分析伪标签无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。