arXiv:2501.15579cs.CVcs.CL2025-01被引 7

首个可解释的医学影像大模型,兼顾诊断精度与医生可理解性。

An Explainable Biomedical Foundation Model via Large-Scale Concept-Enhanced Vision-Language Pre-training

  • 用2300万张图像-文本-概念三元组训练,融合医学术语系统
  • 在52项临床任务中超越现有模型,准确率领先且解释清晰
  • 适合需要可解释AI的医疗场景,如放射科辅助诊断

临床应用人工智能于医学影像需兼具诊断准确性和可解释性。当前多模态医学基础模型侧重性能,其黑箱特性难以解释决策过程。本文提出ConceptCLIP,首个兼具顶尖诊断准确率与人类可理解解释的可解释医学基础模型。构建包含2300万图像-文本-概念三元组的MedConcept-23M数据集,涵盖10种医学影像模态,临床概念源自统一医学语言系统。通过创新的双对齐方法,同时学习全局图像-文本表示与细粒度区域-概念关联,实现精准可解释分析。建立覆盖52个临床任务的最全面评估基准。实验表明,ConceptCLIP显著优于现有先进模型,且临床专家验证其解释具备可读性。作为首个兼具精确性与可解释性的医学基础模型,ConceptCLIP推动可信AI在医学中的广泛应用。

原文摘要 · Abstract (English)

The clinical adoption of artificial intelligence (AI) in medical imaging requires models that are both diagnostically accurate and interpretable to clinicians. While current multimodal biomedical foundation models prioritize performance, their black-box nature hinders explaining the decision-making process in clinically meaningful concepts. Here, we present ConceptCLIP, the first explainable biomedical foundation model that achieves state-of-the-art diagnostic accuracy while delivering human-interpretable explanations across diverse imaging modalities. We curate MedConcept-23M, the largest pre-training dataset comprising 23 million image-text-concept triplets across diverse medical modalities, where clinical concepts are derived from the Unified Medical Language System. Leveraging this dataset, we develop ConceptCLIP through a novel dual-alignment approach that simultaneously learns global image-text representations and fine-grained region-concept associations for precise and interpretable medical image analysis. We curate the most extensive evaluation benchmark for multimodal biomedical foundation models, covering 52 clinical tasks spanning 10 imaging modalities. Extensive experiments demonstrate that ConceptCLIP outperforms existing state-of-the-art multimodal biomedical foundation models. Importantly, ConceptCLIP demonstrates superior diagnostic performance while providing human-understandable explanations validated by clinical experts. As the first precise and interpretable biomedical foundation model, ConceptCLIP represents a critical milestone toward the widespread clinical adoption of AI, thereby advancing trustworthy AI in medicine.

医学AI可解释性视觉语言模型基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。