arXiv:2603.08921cs.CVcs.LG2026-03

将临床指南融入视觉语言模型,提升医学影像推理的可解释性与准确性。

Vision-Language Models Encode Clinical Guidelines for Concept-Based Medical Reasoning

  • 融合临床指南与多模态模型,用结构化概念指导诊断推理。
  • 在超声和乳腺钼靶上分别达到94.2%和84.0%的AUROC,表现优异。
  • 适合医疗AI可解释性研究者及临床辅助决策系统开发者。

概念瓶颈模型(CBMs)通过将视觉特征映射到有意义的概念来实现可解释的AI,其序列结构有助于连接预测与支持概念。在医学影像领域,透明性至关重要,但传统离散概念常忽略诊断指南和专家经验,影响复杂病例的可靠性。本文提出MedCBR,一种结合临床指南的基于概念的推理框架。将标注的临床描述转化为符合指南的文本,并通过多任务目标训练模型,联合对齐图像特征、概念与病理,包括多模态对比对齐、概念监督和诊断分类。随后,推理模型将预测结果转化为结构化临床叙述,模拟基于指南的专家推理。在超声和乳腺钼靶数据集上,诊断的AUROC分别为94.2%和84.0%;非医疗数据集上准确率达86.1%。该框架提升了可解释性,实现了从医学图像分析到决策的端到端闭环。

原文摘要 · Abstract (English)

Concept Bottleneck Models (CBMs) are a prominent framework for interpretable AI that map learned visual features to a set of meaningful concepts for task-specific downstream predictions. Their sequential structure enhances transparency by connecting model predictions to the underlying concepts that support them. In medical imaging, where transparency is essential, CBMs offer an appealing foundation for explainable model design. However, discrete concept representations often overlook broader clinical context such as diagnostic guidelines and expert heuristics, reducing reliability in complex cases. We propose MedCBR, a concept-based reasoning framework that integrates clinical guidelines with vision-language and reasoning models. Labeled clinical descriptors are transformed into guideline-conformant text, and a concept-based model is trained with a multitask objective combining multimodal contrastive alignment, concept supervision, and diagnostic classification to jointly ground image features, concepts, and pathology. A reasoning model then converts these predictions into structured clinical narratives that explain the diagnosis, emulating expert reasoning based on established guidelines. MedCBR achieves superior diagnostic and concept-level performance, with AUROCs of 94.2% on ultrasound and 84.0% on mammography. Further experiments on non-medical datasets achieve 86.1% accuracy. Our framework enhances interpretability and forms an end-to-end bridge from medical image analysis to decision-making.

医学影像可解释AI临床指南多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。