arXiv:2604.09841cs.CVcs.AI2026-04

医学视觉语言模型在复杂任务中表现脆弱,微调未必有效。

Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models

  • 用描述生成+纯文本诊断的流水线测试模型深层知识
  • 高难度医学任务下准确率接近随机,性能随难度下降
  • 适合关注医疗AI可靠性与提示敏感性的研究者

视觉语言模型(VLMs)正被广泛用于医学领域微调,但其是否真正提升临床推理能力仍不明确。本文评估了四组开源模型(LLaVA vs. LLaVA-Med;Gemma vs. MedGemma)在四项医学影像任务上的表现:脑肿瘤、肺炎、皮肤癌和组织病理学分类。随着任务难度增加,模型性能显著下降至接近随机水平,表明其临床推理能力有限。医学微调未带来一致优势,且模型对提示形式极为敏感,微小变化导致准确率和拒答率剧烈波动。为检验封闭式问答是否抑制潜在知识,我们引入基于描述的诊断流程:由模型生成图像描述,再交由纯文本模型(GPT-5.1)诊断。该方法虽恢复部分额外信号,但仍受限于任务难度。视觉编码器嵌入分析显示,失败源于弱视觉表征与下游推理双重问题。总体而言,医学VLM表现脆弱、依赖提示,且无法通过领域微调可靠提升。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly adapted through domain-specific fine-tuning, yet it remains unclear whether this improves reasoning beyond superficial visual cues, particularly in high-stakes domains like medicine. We evaluate four paired open-source VLMs (LLaVA vs. LLaVA-Med; Gemma vs. MedGemma) across four medical imaging tasks of increasing difficulty: brain tumor, pneumonia, skin cancer, and histopathology classification. We find that performance degrades toward near-random levels as task difficulty increases, indicating limited clinical reasoning. Medical fine-tuning provides no consistent advantage, and models are highly sensitive to prompt formulation, with minor changes causing large swings in accuracy and refusal rates. To test whether closed-form VQA suppresses latent knowledge, we introduce a description-based pipeline where models generate image descriptions that a text-only model (GPT-5.1) uses for diagnosis. This recovers a limited additional signal but remains bounded by task difficulty. Analysis of vision encoder embeddings further shows that failures stem from both weak visual representations and downstream reasoning. Overall, medical VLM performance is fragile, prompt-dependent, and not reliably improved by domain-specific fine-tuning.

视觉语言模型医学AI提示敏感性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。