arXiv:2607.04526cs.SDcs.AI2026-07

无需训练,用后处理方法解决异常声音检测的域不平衡问题。

Training-Free Model Selection and Domain-Aware Score Calibration for First-Shot Anomalous Sound Detection

  • 基于冻结音频嵌入的无训练后处理,按域校准阈值。
  • 在DCASE 2025上,预测得分相关性达0.91,提升评估分数至59.34。
  • 适合追求高鲁棒性的工业异常检测应用者。

DCASE Challenge Task 2要求在仅用一个阈值的情况下检测未见过的机器类型异常声音,且无法区分测试片段来自数据丰富的源域(990个正常样本)还是数据稀缺的目标域(10个样本)。组织方指出两个开放问题:不同系统在源域与目标域上的AUC呈负相关,且开发集表现无法预测评估集表现。本文提出一种无需训练的后处理方法,对冻结的音频嵌入进行改进:(i) 基于先验强度m,对各域的分位数校准向合并分布收缩,描绘源/目标平衡边界;(ii) 采用无标签交叉验证的域平衡判据,仅使用正常训练样本对候选配置排序,并结合粗粒度标注可行性过滤。在DCASE 2025中,该判据在45种配置中预测官方评分的斯皮尔曼相关系数为+0.91(95%置信区间[+0.83, +0.95]),而开发集得分相关性仅为+0.06。基于判据的选择将评估分数从55.83提升至59.34(刀切法置信区间[2.2, 4.8]),在扩展网格上进一步达到61.05——回溯排名第三。在2023和2024年复制实验显示,开发集得分始终无效,退化配置重复出现(均被过滤),但仅在2025年判据具备预测力;另两年固定全等化默认方案表现不劣于判据选择。2026年前瞻测试已冻结,所有核心结果经官方评估器复现。

原文摘要 · Abstract (English)

First-shot anomalous sound detection in DCASE Challenge Task 2 must flag anomalies of unseen machine types with a single threshold, without knowing whether a test clip comes from the data-rich source domain (990 normal training clips) or the data-scarce target domain (10). Two organizer-reported problems remain open: source- and target-domain AUC are negatively correlated across systems, and development-set performance does not predict evaluation-set performance. We address both with a training-free post-hoc layer over frozen audio embeddings: (i) per-domain quantile calibration shrunk toward a pooled map by a prior strength m, tracing a source/target balance frontier, and (ii) a label-free cross-validated domain-balance criterion that ranks candidate configurations from training normals only, paired with a coarse development-labeled viability veto. On DCASE 2025, the criterion rank-predicts the official evaluation score across a 45-configuration grid (Spearman rho = +0.91; family-block bootstrap 95% CI [+0.83, +0.95]) while development score is uninformative (+0.06). Criterion-based selection raises the evaluation score from 55.83 to 59.34 (jackknife CI [2.2, 4.8]) and, on an extended grid, to 61.05 -- retrospectively fourth of 35 teams. Replicating on DCASE 2023 and 2024 bounds the claim: development score is uninformative in all three years and degenerate configurations recur (vetoed every time), but under family-clustered uncertainty the criterion's predictive evidence survives only in 2025; in both replication years a fixed full-equalization default matches or beats criterion-based selection. A DCASE 2026 forward test is frozen before the 2026 evaluation ground truth is released; all headline numbers are reproduced by the official evaluator.

异常检测无训练域适应声音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。