arXiv:2608.17370cs.LG2026-08

用最优传输生成可解释临床数据,发现热力图常不能真实定位病灶。

Pathology Transport: Optimal-Transport Explanations for Clinical Data, and When Their Heatmaps (Fail to) Localize Disease

论文配图:Pathology Transport: Optimal-Transport Explanations for Clinical Data, and When Their Heatmaps (Fail to) Localize Disease
图 1 · 摘自论文原文
  • 构建基于最优传输的生成模型,从健康与患病群体分布中提取解释。
  • 在乳腺癌数据上实现0.91的恶性度判别性能,热力图与监督模型相关性达0.5。
  • 在胸片中发现热力图仅反映群体趋势,无法精确定位真实病灶。

生成模型为临床AI可解释性提供新路径:不依赖分类器,而是建模健康与患病人群的分布,并从中读取其几何差异作为解释。我们构建了一个在两个临床分布间训练的最优传输修正流系统,提出一个领域较少检验的核心问题:生成的解释热力图是否真正定位疾病?在乳腺癌肿瘤标志物(乳腺癌威斯康星数据集)上,单一模型生成个体反事实样本,提供无监督恶性度评分(AUROC 0.91;五次种子平均0.93±0.01),以及与监督分类器一致的标签无关归因(相关系数r ~ 0.5),构成紧凑且诚实的可解释引擎,尽管未超越逻辑回归预测性能。转向胸部X光片时,我们发现传输热力图为群体层面信号,非局部定位;基于重建的身份保持变体能定位合成病灶(点位游戏得分0.52),但在真实RSNA放射科医生标注框上退化至随机水平,唯有监督梯度类激活映射(Grad-CAM)保持高于随机。核心发现是:合成到真实间的差距——对植入病灶看起来可信的无标签热力图,并非真实定位的证据。我们贡献了一个可复用的最优传输解释方法及可控基准,用于压力测试其定位能力。

原文摘要 · Abstract (English)

Generative models promise a route to explainable clinical AI: rather than probe a classifier, model the distributions of healthy and diseased patients and read explanations off the geometry between them. We build such a system - an optimal-transport rectified flow trained between two clinical distributions - and use it to ask a pointed question the field too rarely tests: do the resulting explanation heatmaps actually localize disease? On tabular tumour biomarkers (Breast Cancer Wisconsin) a single flow yields per-patient counterfactuals, an unsupervised malignancy score (AUROC 0.91; 0.93 +/- 0.01 across five seeds), and a label-free attribution that agrees with a supervised classifier (r ~ 0.5) - a compact, honest interpretability engine, though it never out-predicts logistic regression. Moving to chest X-rays, we show the transport heatmap is a population-level signal, not a localiser; a reconstruction-based, identity-preserving variant does localize synthetic lesions (pointing game 0.52), yet on real RSNA radiologist boxes it collapses to chance while only supervised Grad-CAM stays above it. The central result is a synthetic-to-real gap: label-free heatmaps that look compelling on planted lesions are not evidence of real localisation. We contribute a reusable optimal-transport recipe for generative explanations and a controlled benchmark for stress-testing whether they localize.

可解释性最优传输临床AI热力图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。