arXiv:2605.14255cs.LGcs.CV2026-05

提出可评估视觉检测模型解释可信度的新方法,发现解释效果取决于模型读出结构。

Architecture-Aware Explanation Auditing for Industrial Visual Inspection

论文配图:Architecture-Aware Explanation Auditing for Industrial Visual Inspection
图 1 · 摘自论文原文
  • 基于读出机制距离设计架构感知的解释审计协议
  • ViT-Tiny+Attention Rollout在删除AUC仅0.211,远低于其他模型
  • 强调解释路径需与模型架构协同设计,适合工业质检领域研究者

工业视觉检测系统日益依赖深度分类器,其热力图解释虽外观合理,却未必反映真实决策依据。本文提出基于原生读出假设的架构感知解释审计协议:解释方法的扰动忠实性受其结构与模型原生决策机制距离限制。在9类、17.2万张的WM-811K晶圆图上,采用三种子零填充扰动协议,ViT-Tiny + Attention Rollout的删除AUC仅为0.211,而Swin-Tiny / ResNet18+CBAM / DenseNet121 + Grad-CAM达到0.432–0.525(绝对Cohen's d > 1.1),尽管前者分类准确率更低。Swin-Tiny显示,尽管是Transformer,其空间特征层次结构使其兼容Grad-CAM,说明决定因素是读出结构而非架构族。一种模型无关控制方法RISE将所有家族的删除AUC压缩至约0.1,表明差距源于解释路径;值得注意的是,RISE优于所有原生方法,说明原生读出是兼容性原则而非最优保证。模糊填充敏感性分析显示,在不同扰动基线下,家族排序反转,强化了忠实性排名是(模型, 解释器, 扰动算子)三元组的联合属性。对MVTec AD(预训练模型)的探索性边界条件研究表明,审计结果依赖数据集/任务,并识别出需特别注意的条件。该协议提供可操作指导:解释路径应基于读出结构与模型架构共同设计,部署热图应附带定量忠实性指标。

原文摘要 · Abstract (English)

Industrial visual inspection systems increasingly rely on deep classifiers whose heatmap explanations may appear visually plausible while failing to identify the image regions that actually drive model decisions. This paper operationalizes an architecture-aware explanation audit protocol grounded in the native-readout hypothesis: the perturbation-based faithfulness of an explanation method is bounded by its structural distance from the model's native decision mechanism. On WM-811K wafer maps (9 classes, 172k images) under a three-seed zero-fill perturbation protocol, ViT-Tiny + Attention Rollout attains Deletion AUC 0.211 against 0.432-0.525 for Swin-Tiny / ResNet18+CBAM / DenseNet121 + Grad-CAM (abs(Cohen's d) > 1.1), despite lower classification accuracy. Swin-Tiny disentangles architecture family from readout structure: despite being a Transformer, its spatial feature-map hierarchy makes it Grad-CAM compatible, showing that the operative factor is readout structure rather than architecture family. A model-agnostic control (RISE) compresses all families to Deletion AUC about 0.1, indicating the gap arises from the explainer pathway; notably, RISE outperforms all native methods, so native readout is a compatibility principle rather than an optimality guarantee. A blur-fill sensitivity analysis shows that the family ordering reverses under a different perturbation baseline, reinforcing that faithfulness rankings are joint properties of (model, explainer, perturbation operator) triples. An exploratory boundary-condition study on MVTec AD (pretrained models) indicates that audit results are dataset/task dependent and identifies conditions requiring qualification. The protocol yields actionable guidance: explanation pathways should be co-designed with model architectures based on readout structure, and deployed heatmaps should be accompanied by quantitative faithfulness metrics.

视觉检测模型解释工业质检忠实性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。