arXiv:2608.25251cs.CVcs.AI2026-08

医学视觉语言模型在分布变化下会失效,因迁移、对齐和代理信息残留。

What Do Medical Vision-Language Models Learn in Radiology? Transfer, Alignment, and Source-Proxy Leakage Under Distribution Shift

论文配图:What Do Medical Vision-Language Models Learn in Radiology? Transfer, Alignment, and Source-Proxy Leakage Under Distribution Shift
图 1 · 摘自论文原文
  • 对比自监督与图像网初始化,发现前者更利于跨数据集迁移。
  • 外部数据测试中配对检索率低,且源数据代理信息仍可被恢复。
  • 模型看似可靠,实则存在隐蔽的迁移与对齐失败,需压力测试验证。

医学视觉语言模型(VLMs)在特定数据集上表现可靠,但在采集域、标注方式或评估协议变化时会失效。本文通过NIH ChestXray14和CheXpert研究源数据到跨数据集的视觉迁移,排除无监督域适应干扰;利用PadChest和OpenI,在严格配对索引检索条件下评估多模态对齐,并量化冻结嵌入中保留的元数据衍生的源代理信息。结果显示,自监督视觉初始化在匹配的ResNet-18结构中优于监督式ImageNet初始化;对抗性适配仅在狭窄范围内有效,随压力增大即失稳。在外部OpenI测试下,多模态精确配对检索率仍较低,源代理信息可从学习表示中恢复。定性分析显示,多数情况具临床合理的跨数据集结构与胸腔注意力模式,但设备相关与假阳性案例仍模糊不清。辅助架构检查结果任务依赖性强,无法支持通用主干网络排序。总体表明,单一协议下的表面能力可能掩盖迁移、对齐及捷径学习缺陷,亟需在分布偏移下进行压力测试评估医学VLM。

原文摘要 · Abstract (English)

Medical vision-language models (VLMs) can appear reliable in-domain while failing when acquisition domain, paired supervision, or evaluation protocol changes. We study this failure mode as a representation-level blind spot relevant to epistemic intelligence, without claiming a formal estimator of epistemic uncertainty. Using NIH ChestXray14 and CheXpert, we first isolate source-only cross-dataset visual transfer from unsupervised domain-adaptation diagnostics. Using PadChest and OpenI, we then evaluate multimodal alignment under strict pair-index retrieval and quantify metadata-derived source-proxy information retained in frozen embeddings. Self-supervised visual initialization improves NIH-to-CheXpert transfer over supervised ImageNet initialization in matched ResNet-18 comparisons, whereas adversarial adaptation is useful only in a narrow regime and becomes unstable as adversarial pressure increases. Multimodal exact-pair retrieval remains low under external OpenI stress testing, and source-proxy information remains recoverable from learned representations. Qualitative nearest-neighbor and Grad-CAM analyses show clinically plausible cross-dataset structure and thoracic attention patterns in many cases, while device-heavy and false-positive cases remain ambiguous. Auxiliary architecture checks are task-dependent and do not support a universal backbone ranking. Overall, the study shows that apparent competence under a single protocol can conceal transfer, alignment, and shortcut-related failure modes, motivating stress-tested evaluation of medical VLMs under distribution shift.

医学视觉语言模型分布外泛化迁移学习多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。