arXiv:2606.10066cs.CVcs.AI2026-06

检测医学视觉语言模型训练数据泄露,发现图像和文本存在分布重叠但非完全记忆。

A Controlled Audit of Pretraining Contamination in Public Medical Vision-Language Benchmarks

论文配图:A Controlled Audit of Pretraining Contamination in Public Medical Vision-Language Benchmarks
图 1 · 摘自论文原文
  • 用四种检测方法分析公开医疗多模态数据集中的预训练污染
  • 19.8%图像在特定模型下被标记为源重叠,文本存在可交换性信号
  • 提醒研究者警惕小规模数据集的误判风险,适合关注模型可信性的研究人员

医学视觉语言模型(VLMs)在公开基准上评估时,常假设其训练数据未包含这些测试样本。本文对SLAKE-En、PathVQA、VQA-RAD及OmniMedVQA镜像数据集进行审计,采用四类检测器:图像侧近邻重叠(针对PMC-OA-beta)、标准顺序可交换性、队列相对最小K%++尾部增强、跨模型前K重叠。结果发现,在SLAKE-En中,使用SigLIP-B-16检测到19.8%图像为源重叠,SigLIP-SO400M为4.2%,而域外控制组无误报(0/2000)。人工判断显示为同模态、同投影但不同患者匹配,属分布重叠而非像素级复制。文本方面,Qwen2.5-VL在SLAKE-En中仍显标准顺序可交换性信号,且不随排序扰动消失,也存在于外部非医疗基线。在OmniMedVQA镜像中,五种模型均出现该信号,唯BLIP-2保持干净。然而,队列相对检测与跨模型重叠在外部预领域基线中失效,即使未接触医疗数据也重现阳性信号,表明此类检测器在小医疗模型队列中不可靠。结论:当前检测方法无法独立验证成员推理,需谨慎解读。

原文摘要 · Abstract (English)

Medical vision-language models (VLMs) are evaluated on public benchmarks whose images and question-answer pairs have been freely downloadable for years, yet reported accuracy assumes these examples were absent from pretraining. We audit open VLMs on SLAKE-En, PathVQA, VQA-RAD, and an auxiliary public OmniMedVQA mirror using four detector families: image-side near-neighbour overlap against PMC-OA-beta, canonical-order exchangeability, cohort-relative Min-K%++ tail enrichment, and cross-model top-K overlap. We find measurable image-side source overlap on SLAKE-En: 19.8% of images are flagged under SigLIP-B-16 and 4.2% under SigLIP-SO400M, while out-of-domain controls produce 0/2000 flags. Manual adjudication shows same-modality, same-projection matches to different patients rather than verified pixel-level duplicates, so we interpret this as source or distributional overlap rather than confirmed per-image memorization. On the text side, Qwen2.5-VL on SLAKE-En shows a canonical-order exchangeability signal that survives ordering ablation and external non-medical baselines. On the OmniMedVQA mirror, exchangeability fires for five medical and general VLMs while BLIP-2 remains clean. In contrast, cohort-relative Min-K%++ tail enrichment and cross-model top-K overlap collapse under an external pre-domain baseline: BLIP-2 reproduces the apparent positive signals despite lacking plausible medical-VQA exposure. We conclude that these cohort-relative detectors are unreliable as standalone membership-inference signals on small medical-VLM cohorts.

医学多模态模型审计数据泄露成员推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。