arXiv:2604.11868cs.CV2026-04

无监督发现医学视觉模型中的潜在概念,提升可解释性。

MedConcept: Unsupervised Concept Discovery for Interpretability in Medical VLMs

论文配图:MedConcept: Unsupervised Concept Discovery for Interpretability in Medical VLMs
图 1 · 摘自论文原文
  • 从预训练模型中无监督提取稀疏神经激活,生成类报告摘要。
  • 通过独立医学大模型评估,量化概念与影像报告的语义一致性。
  • 适合希望理解医学AI决策过程的研究者和临床医生。

尽管医学视觉-语言模型(VLM)在肿瘤或器官分割、诊断预测等任务上表现优异,其隐含表示的不透明性限制了临床信任和解释能力。当前可解释性方法如梯度或注意力可视化多局限于分类任务,且无法提供可跨下游任务复用的概念级解释。本文提出MedConcept,一种完全无监督的方法,用于发现预训练医学VLM中的潜在医学概念,并将其与临床可验证的文本语义对齐。该方法从预训练表示中识别稀疏神经元级概念激活,并转换为类报告风格的摘要,支持医师级推理过程检查。为解决概念解释缺乏定量评估的问题,引入基于独立预训练医学LLM的语义验证协议,定义三类概念得分:对齐(Aligned)、不一致(Unaligned)和不确定(Uncertain),分别衡量概念与放射科报告的语义支持、矛盾或模糊性,仅用于事后评估。这些得分提供了医学VLM可解释性的量化基准。所有代码、提示和数据将在论文接受后公开。

原文摘要 · Abstract (English)

While medical Vision-Language models (VLMs) achieve strong performance on tasks such as tumor or organ segmentation and diagnosis prediction, their opaque latent representations limit clinical trust and the ability to explain predictions. Interpretability of these multimodal representations are therefore essential for the trustworthy clinical deployment of pretrained medical VLMs. However, current interpretability methods, such as gradient- or attention-based visualizations, are often limited to specific tasks such as classification. Moreover, they do not provide concept-level explanations derived from shared pretrained representations that can be reused across downstream tasks. We introduce MedConcept, a framework that uncovers latent medical concepts in a fully unsupervised manner and grounds them in clinically verifiable textual semantics. MedConcept identifies sparse neuron-level concept activations from pretrained VLM representations and translates them into pseudo-report-style summaries, enabling physician-level inspection of internal model reasoning. To address the lack of quantitative evaluation in concept-based interpretability, we introduce a quantitative semantic verification protocol that leverages an independent pretrained medical LLM as a frozen external evaluator to assess concept alignment with radiology reports. We define three concept scores, Aligned, Unaligned, and Uncertain, to quantify semantic support, contradiction, or ambiguity relative to radiology reports and use them exclusively for post hoc evaluation. These scores provide a quantitative baseline for assessing interpretability in medical VLMs. All codes, prompt and data to be released on acceptance. Ke

可解释性医学AI视觉语言模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。