用轻量探针从冻结模型中提取医学视觉问答信号,效果优于复杂系统。
MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

- 通过线性探针从冻结的视觉语言模型中直接预测答案
- 在多个数据集上比提示工程和医疗微调模型表现更好
- 适合研究模型内部表征或资源受限场景下的医学问答
医学视觉问答(Med-VQA)通常被认为需要医疗微调、大模型或复杂多智能体系统。我们提出轻量级探针框架 MedProb,无需生成自由文本,即可从冻结的视觉语言模型(VLM)表征中预测多项选择题答案。在 PATH-VQA、SLAKE 和 VQA-RAD 上,MedProb 捕获的答案相关信号显著多于提示工程,且优于医疗 VLM 与智能体系统。探针方法还缩小了小模型与大模型之间的性能差距,表明小 VLM 中存在更多可恢复的医学科问答信号。在14组通用与医疗适配的 VLM 对比中,医疗适应并未持续提升线性可解性。自由文本生成存在最高达10个百分点的答案位置偏差,而 MedProb 也有位置偏差,但影响机制不同。主要结果针对多项选择/多分类医学科问答任务;此外,通过拒绝采样评分过程,探针可扩展至开放式生成。
原文摘要 · Abstract (English)
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text generation. Across PATH-VQA, SLAKE, and VQA-RAD, MedProb recovers substantially more answer-relevant signal than prompting and performs stronger than medical VLMs and agentic systems. Probing also reduces the apparent gap between small and large models compared to prompting, suggesting that smaller VLMs contain more recoverable Med-VQA signal than generation-based evaluation reveals. Across 14 matched general-purpose and medical VLM pairs, medical adaptation does not consistently improve this linear decodability. Finally, free-text generation exhibits an answer-position bias of up to 10 percentage points, whereas MedProb also has positional bias, however, it is impacted differently than prompting. Our main results target the multiple-choice/multiclass Med-VQA setting; we additionally show the probe can be extended to open-ended generation via a rejection-sampling scoring procedure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。