让癌症分割模型自己识别错误,用内部概念激活发现潜在失败。
Medical AI Encodes a "Feeling of Error": Verifying Cancer Segmentation via Internal Concepts

- 通过分解神经激活捕捉模型内部的错误信号。
- 失败案例中概念激活数量和强度均显著降低。
- 比传统方法更准地发现错误,且不损害分割质量。
癌症分割模型可能在无声中出错,生成看似合理却错误的掩码,导致漏诊或不必要的活检。一个关键问题是:人工智能是否‘知道’自己错了?若能,能否利用这种信号预测其自身失败?人类拥有‘错误感’(Feeling of Error, FOE)——一种自发的不安感,可标记思维中的潜在错误。本文研究癌症分割模型是否存在类似的内部信号。不同于输出层面的提示(如预测置信度或不确定性),这些方法无法揭示失败原因,且检测敏感性与分割质量存在权衡。本文提出从模型内部机制中捕捉类似错误感的信号。利用机制可解释性工具(如稀疏自编码器),我们将内部神经激活分解为一组可人工理解的概念,并发现失败案例具有独特的潜在特征:活跃概念较少且激活幅度较低。通过训练分类器基于这些概念激活,实现准确的失败检测并提供错误解释。在前列腺、胰腺和脑癌分割任务上的实验表明,该方法在失败检测上优于输出基线,同时保持了分割质量。
原文摘要 · Abstract (English)
Cancer segmentation models can fail silently, generating plausible but incorrect masks that risk missed findings or unnecessary biopsies. A critical question arises: Do AI models "know" when they are wrong, and if so, can we use the signal to predict their own failures? Humans do have a "Feeling of Error" (FOE): a spontaneous sense of unease that flags a potential error during thinking. We investigate whether cancer segmentation models exhibit an analogous internal signal. Unlike output-level cues (e.g., prediction confidence or uncertainty), which offer no insight into why a failure occurs and suffer from a sensitivity-quality tradeoff where high detection sensitivity could degrade overall segmentation quality. We instead propose to capture the model's FOE from its inner workings. Using mechanistic interpretability tools, specifically Sparse Autoencoders, we decompose internal neural activations into a dictionary of human-interpretable concepts and show that failure cases exhibit a distinct latent signature: fewer active concepts with lower activation magnitudes compared to successful segmentation. By training a classifier on these concept activations, we achieve accurate failure detection along with explanations for the model's mistakes. Experiments on prostate, pancreatic, and brain cancer segmentation demonstrate that our approach outperforms output-based methods in failure detection while preserving segmentation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。