arXiv:2509.24231cs.CV2025-09被引 2

医学多模态模型可解释诊断,提升临床信任度

EVLF-FM: Explainable Vision Language Foundation Model for Medicine

  • 融合视觉与语言的统一模型,支持多病种诊断与逐像素解释
  • 内部测试准确率达0.858,跨九模态定位精度mIOU达0.743
  • 具备可追踪推理路径,适合医疗AI落地与可信决策

尽管基础模型在医疗AI中前景广阔,但现有系统仍受限于单一模态且缺乏透明推理过程,阻碍临床应用。为弥补这一差距,我们提出EVLF-FM,一种统一多病种诊断能力与细粒度可解释性的医学多模态视觉语言基础模型(VLM)。模型训练与测试覆盖全球23个数据集共130万+样本,涵盖六大学科(皮肤科、肝病学、眼科、病理学、呼吸科、放射科)的十一种影像模态。外部验证使用来自五个模态的10个额外数据集中的8,884个独立测试样本。技术上,EVLF-FM支持多疾病诊断与视觉问答,具备像素级视觉定位与推理能力。内部验证中,其平均准确率(0.858)和F1分数(0.797)优于主流通用与专科模型;在医学视觉定位任务中,跨九模态平均mIOU达0.743,[email protected]为0.837。外部验证进一步证实其零样本与少量样本下具有竞争力的性能,且模型规模更小。通过结合监督学习与视觉强化微调的混合训练策略,EVLF-FM不仅实现顶尖准确率,还能生成与视觉证据对齐的逐步推理过程。EVLF-FM是首个具备可解释性与推理能力的多病种医学基础模型,有望推动真实世界临床部署中基础模型的采纳与信任。

原文摘要 · Abstract (English)

Despite the promise of foundation models in medical AI, current systems remain limited - they are modality-specific and lack transparent reasoning processes, hindering clinical adoption. To address this gap, we present EVLF-FM, a multimodal vision-language foundation model (VLM) designed to unify broad diagnostic capability with fine-grain explainability. The development and testing of EVLF-FM encompassed over 1.3 million total samples from 23 global datasets across eleven imaging modalities related to six clinical specialties: dermatology, hepatology, ophthalmology, pathology, pulmonology, and radiology. External validation employed 8,884 independent test samples from 10 additional datasets across five imaging modalities. Technically, EVLF-FM is developed to assist with multiple disease diagnosis and visual question answering with pixel-level visual grounding and reasoning capabilities. In internal validation for disease diagnostics, EVLF-FM achieved the highest average accuracy (0.858) and F1-score (0.797), outperforming leading generalist and specialist models. In medical visual grounding, EVLF-FM also achieved stellar performance across nine modalities with average mIOU of 0.743 and [email protected] of 0.837. External validations further confirmed strong zero-shot and few-shot performance, with competitive F1-scores despite a smaller model size. Through a hybrid training strategy combining supervised and visual reinforcement fine-tuning, EVLF-FM not only achieves state-of-the-art accuracy but also exhibits step-by-step reasoning, aligning outputs with visual evidence. EVLF-FM is an early multi-disease VLM model with explainability and reasoning capabilities that could advance adoption of and trust in foundation models for real-world clinical deployment.

医学AI可解释性多模态基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。