arXiv:2506.17337eess.IVcs.AI2025-06被引 2

通用视觉语言模型经微调后可媲美甚至超越专业医疗模型。

Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights

  • 用高效微调让通用模型适应医疗任务
  • 在罕见或未知医学模态上表现更优
  • 适合资源有限但需快速部署的临床场景

视觉语言模型(VLMs)在自动化临床图像诊断方面展现出潜力。然而,开发专用医疗VLM需要大量计算资源和精心构建的数据集,且尚不明确通用模型与专用模型在何种条件下各具优势。本研究揭示了专用医疗模型与通用模型的互补性:专用模型在模态对齐任务中仍具价值,但经过高效微调的通用模型在多数任务中表现相当甚至更优,尤其在跨域迁移至未见或罕见的外部数据分布(OOD)医学模态时。结果表明,通用模型虽缺乏专业医疗预训练,却可能为临床AI发展提供一种可扩展、低成本的路径。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings. However, developing specialist medical VLMs requires substantial computational resources and carefully curated datasets, and it remains unclear under which conditions generalist and specialist medical VLMs each perform best. This study highlights the complementary strengths of specialist medical and generalist VLMs. Specialists remain valuable in modality-aligned use cases, but we find that efficiently fine-tuned generalist VLMs can achieve comparable or even superior performance in most tasks, particularly when transferring to unseen or rare OOD medical modalities. These results suggest that generalist VLMs, rather than being constrained by their lack of specialist medical pretraining, may offer a scalable and cost-effective pathway for advancing clinical AI development.

视觉语言模型医疗AI模型微调跨模态迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。