arXiv:2602.10508cs.CV2026-02被引 1

通过潜空间对比分析,揭示医学图像分割模型的错误根源并修复。

Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation

  • 用稀疏自编码器分解模型激活,提取可解释的潜在特征。
  • 无需重训练即可修复70%错误,Dice分数从39.4%提升至74.2%。
  • 适合关注模型可解释性与跨数据集泛化的研究者。

现代分割模型虽表现强劲,但透明度低,难以诊断失败、理解数据分布偏移或进行有依据的干预。本文提出Med-SegLens,一种基于潜空间对比的模型差分框架,利用在SegFormer和U-Net上训练的稀疏自编码器,将模型激活分解为可解释的潜在特征。通过在健康、成人、儿童及撒哈拉以南非洲胶质瘤队列间进行跨架构与跨数据集的潜空间对齐,我们发现一组稳定的共享表示,而数据分布偏移主要由人群特异性潜变量的差异依赖驱动。实验表明,这些潜变量是分割失败的因果瓶颈,针对性的潜空间干预可纠正错误且无需重训练,在70%的失败案例中恢复性能,使Dice分数从39.4%提升至74.2%。结果证明,潜空间模型差分可作为诊断失败与缓解数据分布偏移的实用且机制化工具。

原文摘要 · Abstract (English)

Modern segmentation models achieve strong predictive performance but remain largely opaque, limiting our ability to diagnose failures, understand dataset shift, or intervene in a principled manner. We introduce \textbf{Med-SegLens}, a model-diffing framework that decomposes segmentation model activations into interpretable latent features using sparse autoencoders trained on SegFormer and U-Net. Through cross-architecture and cross-dataset latent alignment across healthy, adult, pediatric, and sub-Saharan African glioma cohorts, we identify a stable backbone of shared representations, while dataset shift is driven by differential reliance on population-specific latents. We show that these latents act as causal bottlenecks for segmentation failures, and that targeted latent-level interventions can correct errors and improve cross-dataset adaption without retraining, recovering performance in 70% of failure cases and improving Dice score from 39.4% to 74.2%. Our results demonstrate that latent-level model diffing provides a practical and mechanistic tool for diagnosing failures and mitigating dataset shift in segmentation models.

可解释性医学图像潜空间模型差分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。