无需训练,专家直接修改内存库即可提升异常检测效果。
Training-Free Human-in-the-Loop Anomaly Detection via Memory Bank Correction

- 专家通过直接编辑内存库纠正错误,无需重新训练或梯度计算。
- 仅用10个正常样本加人工修正,性能接近上百样本训练的模型。
- 适合工厂现场无工程师时快速部署,尤其对小样本场景优化明显。
异常检测在训练数据最匮乏的场景下最难部署:新产线仅有少量‘黄金’样本,且现场无机器学习工程师。本文提出一种无需训练的人类介入框架,领域专家通过直接编辑PatchCore的内存库进行修正:无需重训练、无需梯度、无需原始训练数据。误报修正时,系统通过自校准的新颖性门控,仅将距离中位数正常样本最近邻超过阈值的图像块加入内存库。基于仅十枚黄金样本构建的内存库,经操作员修正后,平均缩小了66%的性能差距(均值80%),在15个MVTec AD类别中显著提升12个,无任何类别受损;十样本加修正结果优于数百样本未修正的情况。对于已训练内存库,改进空间较小但集中在采样不足的正常外观区域(牙刷+0.10,金属螺母+0.09,拉链+0.05,螺丝+0.05),仅网格类别有轻微下降。评估采用保留测试协议(每类20次划分,Holm校正威尔科克森检验),因修正图像进入内存库会因记忆导致假性AUROC逼近1.0。被动与主动查询无统计差异;控制实验表明收益主要来自部署阶段标签生成,成本仅为全量审查的43%;缺陷记忆扩展方法彻底失败。反馈基于真实标签模拟,实际专家测试中误标代价高,尚待后续研究。
原文摘要 · Abstract (English)
Anomaly detectors are hardest to deploy exactly where training data is scarcest: a newly commissioned production line has a handful of verified "golden" samples and no machine-learning engineer on the factory floor. We present a training-free human-in-the-loop framework in which a domain expert corrects a PatchCore detector by direct memory bank editing: no retraining, no gradients, no original training data. A false-positive correction inserts the reviewed image's normal patches through a self-calibrating novelty gate admitting only those beyond the median pool-normal nearest-neighbour distance. From a bank built on only ten golden samples, operator corrections close a median 66% of the gap to an uncorrected fully trained bank (mean 80%, raised by three categories that overshoot parity), significantly improving 12 of 15 MVTec AD categories and harming none: ten samples plus corrections outperform hundreds of samples without them. On already-trained banks the headroom is smaller and concentrated where the bank undersamples normal appearance (gated: toothbrush +0.10, metal nut +0.09, zipper +0.05, screw +0.05), and no category except grid is significantly harmed. Evaluation uses a held-out protocol (20 splits per category, Holm-corrected Wilcoxon), because corrected images entering the bank inflate naive evaluation toward AUROC 1.0 by memorisation. Passive and active querying are statistically indistinguishable; a matched-label-budget control attributes gains to deployment-time label production at 43% of exhaustive-review cost; a defect-memory extension fails decisively. Feedback is simulated from ground truth; live expert trials, where mislabelling is costliest on small banks, remain future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。