arXiv:2501.02922cs.CVcs.AI2025-01被引 10

用病理概念解释癌变图像,无需人工标注即可精准定位肿瘤区域。

Label-free Concept Based Multiple Instance Learning for Gigapixel Histopathology

  • 通过视觉语言模型自动识别病理概念,替代人工标注
  • 在两个数据集上准确率超0.9,前20个关键区域85%以上位于肿瘤区
  • 生成医生能理解的概念解释,适合医疗AI可解释性研究

多实例学习(MIL)方法可在仅提供整张切片标签的情况下分析吉字节级全切片图像(WSI)。在高风险医疗领域部署此类算法时,可解释性至关重要。传统MIL方法通过热图显示显著区域,但对用户帮助有限。为此,我们提出一种内在可解释的WSI分类新方法,利用人类可理解的病理概念生成解释。所提出的概念多实例学习(Concept MIL)模型借助视觉语言模型,直接基于图像特征预测病理概念。模型预测通过WSI中顶部K个图像块的关键词线性组合获得,实现通过追踪每个概念对预测的影响来实现内在解释。与传统基于概念的可解释模型不同,本方法无需昂贵的人工标注,而是依赖视觉语言模型。我们在两个常用病理数据集Camelyon16和PANDA上验证该方法,在两个数据集上均达到超过0.9的AUC和准确率,达到当前最佳水平。进一步发现,87.1%(Camelyon16)和85.3%(PANDA)的前20个关键图像块位于肿瘤区域。用户研究表明,模型识别出的概念与病理学家使用的概念高度一致,表明其在人机可解释性方面具有巨大潜力。

原文摘要 · Abstract (English)

Multiple Instance Learning (MIL) methods allow for gigapixel Whole-Slide Image (WSI) analysis with only slide-level annotations. Interpretability is crucial for safely deploying such algorithms in high-stakes medical domains. Traditional MIL methods offer explanations by highlighting salient regions. However, such spatial heatmaps provide limited insights for end users. To address this, we propose a novel inherently interpretable WSI-classification approach that uses human-understandable pathology concepts to generate explanations. Our proposed Concept MIL model leverages recent advances in vision-language models to directly predict pathology concepts based on image features. The model's predictions are obtained through a linear combination of the concepts identified on the top-K patches of a WSI, enabling inherent explanations by tracing each concept's influence on the prediction. In contrast to traditional concept-based interpretable models, our approach eliminates the need for costly human annotations by leveraging the vision-language model. We validate our method on two widely used pathology datasets: Camelyon16 and PANDA. On both datasets, Concept MIL achieves AUC and accuracy scores over 0.9, putting it on par with state-of-the-art models. We further find that 87.1\% (Camelyon16) and 85.3\% (PANDA) of the top 20 patches fall within the tumor region. A user study shows that the concepts identified by our model align with the concepts used by pathologists, making it a promising strategy for human-interpretable WSI classification.

病理图像可解释性概念学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。