arXiv:2506.10669cs.CVcs.AI2025-06被引 4

PiPViT用视觉变换器学习可解释的视网膜图像原型,提升诊断透明度。

PiPViT: Patch-based Visual Interpretable Prototypes for Retinal Image Analysis

  • 基于视觉变换器捕捉图像块间长程依赖,生成可读性强的病变原型。
  • 在4个视网膜OCT数据集上性能媲美顶尖方法,且原型与临床标志物高度相关。
  • 适合需要模型可解释性的医学影像诊断场景,助力临床决策理解。

基于原型的方法通过学习细粒度局部原型提升可解释性,但在像素空间的可视化常与人类可理解的生物标志物不一致。现有方法通常生成过于精细的原型,在医学影像中难以解读,而病变的存在与范围均至关重要。为此,我们提出PiPViT(Patch-based Visual Interpretable Prototypes),一种内在可解释的原型模型。利用视觉变换器(ViT)捕获图像块间的长程依赖,仅使用图像级标签即可学习鲁棒且人类可理解的原型,以近似病变范围。同时,结合对比学习和多尺度输入处理,实现跨尺度生物标志物的有效定位。在四个视网膜OCT图像分类数据集上评估,其性能与最先进方法相当,且生成的解释更具有语义和临床意义。定量验证表明,所学原型在独立测试集上具备显著临床相关性。我们相信PiPViT能透明化决策过程,辅助临床理解诊断结果。

原文摘要 · Abstract (English)

Background and Objective: Prototype-based methods improve interpretability by learning fine-grained part-prototypes; however, their visualization in the input pixel space is not always consistent with human-understandable biomarkers. In addition, well-known prototype-based approaches typically learn extremely granular prototypes that are less interpretable in medical imaging, where both the presence and extent of biomarkers and lesions are critical. Methods: To address these challenges, we propose PiPViT (Patch-based Visual Interpretable Prototypes), an inherently interpretable prototypical model for image recognition. Leveraging a vision transformer (ViT), PiPViT captures long-range dependencies among patches to learn robust, human-interpretable prototypes that approximate lesion extent only using image-level labels. Additionally, PiPViT benefits from contrastive learning and multi-resolution input processing, which enables effective localization of biomarkers across scales. Results: We evaluated PiPViT on retinal OCT image classification across four datasets, where it achieved competitive quantitative performance compared to state-of-the-art methods while delivering more meaningful explanations. Moreover, quantitative evaluation on a hold-out test set confirms that the learned prototypes are semantically and clinically relevant. We believe PiPViT can transparently explain its decisions and assist clinicians in understanding diagnostic outcomes. Github page: https://github.com/marziehoghbaie/PiPViT

可解释性视网膜图像视觉Transformer原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。