arXiv:2601.14774cs.CV2026-01中稿 · The IEEE Internati…

医学视觉语言模型的专科化能提升病变判别力吗?

Does medical specialization of VLMs enhance discriminative power?: A comprehensive investigation through feature distribution analysis

  • 通过特征分布分析对比医疗与非医疗VLM的表示能力
  • 医疗VLM在多模态病变分类任务中展现有效判别特征
  • 文本编码器优化比大量医疗图像训练更重要

本研究分析了公开可用开源医学视觉语言模型(VLM)产生的特征表示。尽管医学VLM应捕捉诊断相关特征,但其学习到的表示仍缺乏深入探索,标准评估如分类准确率无法充分揭示其是否具备真正判别性、病变特异性的特征。理解这些表示对揭示医学图像结构及提升下游任务性能至关重要。本研究旨在探究医学VLM所学特征分布,并评估医学专业化的影响。我们在多个模态的病变分类数据集上,分析代表性医学VLM在多模态图像中提取的特征分布,并与非医学VLM进行比较,以评估领域特定训练效果。实验表明,医学VLM能够提取对医学分类任务有效的判别性特征。此外,近期通过上下文增强改进的非医学VLM(如LLM2CLIP)产生更精细的特征表示。结果表明,提升文本编码器比在医疗图像上密集训练更为关键。值得注意的是,非医学模型对图像上叠加文本字符串引入的偏差尤为敏感。这些发现强调,在选择模型时需根据下游任务谨慎考量,避免因背景偏差(如图像中的文本信息)带来的推理风险。

原文摘要 · Abstract (English)

This study investigates the feature representations produced by publicly available open source medical vision-language models (VLMs). While medical VLMs are expected to capture diagnostically relevant features, their learned representations remain underexplored, and standard evaluations like classification accuracy do not fully reveal if they acquire truly discriminative, lesion-specific features. Understanding these representations is crucial for revealing medical image structures and improving downstream tasks in medical image analysis. This study aims to investigate the feature distributions learned by medical VLMs and evaluate the impact of medical specialization. We analyze the feature distribution of multiple image modalities extracted by some representative medical VLMs across lesion classification datasets on multiple modalities. These distributions were compared them with non-medical VLMs to assess the domain-specific medical training. Our experiments showed that medical VLMs can extract discriminative features that are effective for medical classification tasks. Moreover, it was found that non-medical VLMs with recent improvement with contextual enrichment such as LLM2CLIP produce more refined feature representations. Our results imply that enhancing text encoder is more crucial than training intensively on medical images when developing medical VLMs. Notably, non-medical models are particularly vulnerable to biases introduced by overlaied text strings on images. These findings underscore the need for careful consideration on model selection according to downstream tasks besides potential risks in inference due to background biases such as textual information in images.

医学AI视觉语言模型特征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。