用统一融合公式结合视觉与3D证据,提升工业缺陷检测精度
TC-MAF: Train-Calibrated Bounded Multi-Evidence Fusion for Multimodal Industrial Anomaly Detection

- 固定像素级融合结构,整合多模态检测与一致性线索
- 图像级AUROC达0.979,像素级AUPRO达0.990,性能领先
- 无需测试时调参,适用于少样本与缺失3D数据场景
多模态异常检测得益于互补的RGB与3D信息,但辅助的RGB重建可靠性在不同产品类别间差异大,且通常缺乏类别级测试时策略选择。我们提出TC-MAF,一种基于基础锚点的多证据融合设计,通过单一固定像素级融合公式,整合多模态检测器、互补的异常证据以及轻量级跨模态一致性线索。引入轻量级训练分散置信度(TDC)项,仅使用正常样本训练统计量对辅助模块参与度进行缩放。在MVTec-3D上,TC-MAF实现0.979图像级AUROC与0.990像素级AUPRO,多项式指标均优于对比方法。系统性消融实验表明,融合结构本身是主导因素,而TDC带来可复现的小幅校准增益。额外实验显示该设计在汇总统计变体、辅助分支与主干替换、少样本设置、缺失3D数据及跨数据集评估(Eyecandies)下仍有效。代码已开源。
原文摘要 · Abstract (English)
Multimodal anomaly detection benefits from complementary RGB and 3D evidence, yet auxiliary RGB reconstruction is not equally reliable across product categories and class-wise test-time policy selection is usually unavailable. We propose TC-MAF, a base-anchored multi-evidence fusion design that combines a multimodal detector, complementary Dinomaly evidence, and a small cross-modal consistency cue under one fixed pixel-level fusion formula. A lightweight training-dispersion confidence (TDC) term scales auxiliary participation using only normal training statistics. On MVTec-3D, TC-MAF reaches 0.979 image-level AUROC and 0.990 pixel-level AUPRO, achieving the best mean results on both detection and localization among the compared multimodal methods. Systematic ablations show that the fusion structure itself is the dominant factor, while TDC provides a smaller but reproducible calibration gain over no calibration or arbitrary calibration. Additional experiments show that the same design remains effective under a pooled-statistics variant, auxiliary-branch and backbone substitutions, few-shot settings, a missing-3D setting, and cross-dataset evaluation on Eyecandies. Code is available at https://anonymous.4open.science/r/TC_MAF-C3BB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。