arXiv:2607.19372eess.IVq-bio.QM2026-07

提出可复现的评估框架,检验肠镜分类器得分与定位是否一致。

endoExplain: A reproducible protocol for auditing score-localisation discordance in colonoscopy image classifiers

论文配图:endoExplain: A reproducible protocol for auditing score-localisation discordance in colonoscopy image classifiers
图 1 · 摘自论文原文
  • 构建统一测试流程,对比多种注意力图生成方法
  • 发现62.2%情况下热点不在病灶内,且结果依赖算法和模型
  • 适用于医疗AI可信性审查,尤其关注模型推理可靠性

高分类得分与合理类激活图(CAM)常被同时呈现,但二者均不能单独证明对方可靠。本文提出endoExplain,一个可复现的审计协议,用于检测分类得分与定位之间的不一致,而非新检测器或解释算法。通过内容哈希将HyperKvasir训练集与1000张掩码图像分离后训练EfficientNet-B0、ResNet-34和ConvNeXt-Tiny各三组随机种子。仅使用验证数据进行温度缩放校准,得分从0.0167降至0.0115。在保留的172张得分≥0.90图像上,不同CAM方法的峰值位于病灶外的比例为4.1%至62.2%。空间对齐与删除响应不可互换,随机删除也导致正向激活下降,削弱仅凭删除法得出的特异性结论。外部掩码集结果受数据集影响显著。源类别审计发现155个阳性样本中149个为染色抬举息肉,限制了分类器声明的适用范围。endoExplain使校准、空间一致性、扰动响应与迁移能力可分别检验,提醒勿以单一得分或视觉吸引的CAM作为模型定位或推理依据。

原文摘要 · Abstract (English)

Background and objective: A high classifier score and a plausible class-activation map (CAM) are often presented together, although neither establishes that the other is reliable. We introduce endoExplain as a reproducible protocol for auditing score-localisation discordance rather than as a new detector or explanation algorithm. Methods: Content hashing separated HyperKvasir development images from 1,000 masked images before training. EfficientNet-B0, ResNet-34 and ConvNeXt-Tiny were trained with three seeds each. Scores were temperature scaled using validation data only. Grad-CAM, Grad-CAM++, XGrad-CAM, HiResCAM and Eigen-CAM were evaluated on identical image-mask pairs, alongside random and centre baselines. Outcomes combined peak localisation, overlap, a top-20% deletion response, score-threshold sensitivity and adjustment for lesion size and centrality. The selected checkpoint was transferred without retraining to three external mask cohorts. Results: Temperature scaling reduced test expected calibration error from 0.0167 to 0.0115. Among 172 reserved images with scaled score at least 0.90, peak-outside-lesion rates ranged from 4.1% to 62.2% across CAMs. Method dependence remained evident across architectures and seeds, although method rankings were not universal. Spatial alignment and deletion response were not interchangeable. A random-deletion control also produced positive logit drops, limiting specificity claims based on deletion alone. External positive-mask results were dataset dependent. A source-category audit also exposed that 149/155 test positives were dyed-lifted polyps, materially bounding classifier claims. Conclusions: endoExplain makes calibration, spatial agreement, perturbation response and transfer separately inspectable. The results caution against using a score or visually persuasive CAM as evidence of lesion localisation or model reasoning.

医疗AI模型可解释性肠镜分析可信度审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。