arXiv:2511.08402cs.CVcs.AI2025-11中稿 · Winter Conference …被引 6

医学影像诊断新模型,能精准定位关键解剖结构并结合临床知识

Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation

  • 分层级提取影像中的解剖特征,实现细粒度分析
  • 在跨分布数据集上表现优异,零样本下也能准确解读
  • 适合医疗AI研发者与放射科医生参考

由于影像异质性,医学影像的疾病解析仍具挑战。要达到专家级诊断,需融合细微图像特征与临床知识。现有主流视觉-语言模型将图像视为整体,忽略对诊断至关重要的细粒度信息。医生通过先验医学知识,将解剖结构作为关注区域(ROIs)进行分析。受此人类诊疗流程启发,我们提出Anatomy-VLM,一种具备多尺度信息整合能力的细粒度视觉-语言模型。首先,设计模型编码器从全图中定位关键解剖特征;其次,为这些区域注入结构化知识以实现上下文感知的解释;最后,对多尺度医学信息进行对齐,生成可临床解释的疾病预测。Anatomy-VLM在分布内与分布外数据集上均表现卓越,并在下游图像分割任务中验证了其对解剖与病理性知识的捕捉能力。此外,该模型编码器支持零样本解剖层面解读,展现出强大的专家级临床推理能力。

原文摘要 · Abstract (English)

Accurate disease interpretation from radiology remains challenging due to imaging heterogeneity. Achieving expert-level diagnostic decisions requires integration of subtle image features with clinical knowledge. Yet major vision-language models (VLMs) treat images as holistic entities and overlook fine-grained image details that are vital for disease diagnosis. Clinicians analyze images by utilizing their prior medical knowledge and identify anatomical structures as important region of interests (ROIs). Inspired from this human-centric workflow, we introduce Anatomy-VLM, a fine-grained, vision-language model that incorporates multi-scale information. First, we design a model encoder to localize key anatomical features from entire medical images. Second, these regions are enriched with structured knowledge for contextually-aware interpretation. Finally, the model encoder aligns multi-scale medical information to generate clinically-interpretable disease prediction. Anatomy-VLM achieves outstanding performance on both in- and out-of-distribution datasets. We also validate the performance of Anatomy-VLM on downstream image segmentation tasks, suggesting that its fine-grained alignment captures anatomical and pathology-related knowledge. Furthermore, the Anatomy-VLM's encoder facilitates zero-shot anatomy-wise interpretation, providing its strong expert-level clinical interpretation capabilities.

医学影像细粒度分析视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。