arXiv:2607.12278cs.CVcs.LG2026-07

发现病理图像模型评测存在严重数据泄露,导致性能虚高。

Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

论文配图:Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks
图 1 · 摘自论文原文
  • 通过追踪病例和机构标识,发现训练测试集重叠高达92.3%~100%
  • 模型在泄露数据上的准确率远高于清理后数据,差距明显
  • 适合关注医学AI评测可信度的研究者和审稿人

近期用于计算病理学的视觉语言模型(VLMs)在全切片图像(WSI)视觉问答(VQA)基准上展现出惊人的零样本性能。我们对此进行审计,发现其性能被两个层级的数据泄露严重削弱:患者级泄露(同一病例的切片同时出现在训练和测试集中),以及机构级泄露(不同病例因共享染色批次和扫描仪特征而产生关联)。通过追踪主要公开资源中的切片、病例和组织来源站点(TSS)标识,我们发现基于TCGA的基准中病例级训练测试重叠达92.3%~100%,且几乎完全的TSS重叠。进一步表明,这两种泄露均可在线性可分的基座模型特征空间中被识别,且在已发布检查点上,泄露案例与审计清洁案例间存在显著准确率差异。多个已发表的WSI VLM模型中,最高报告准确率均集中在污染最严重的基准上。因此,当前的WSI VQA评估无法区分真正的多模态推理与对记忆化机构及患者特异性特征的最近邻检索。最后,我们提出具体建议以实现无污染评估,包括改进基准构建、披露数据来源和自动化重叠审计,旨在引导未来研究做出可验证的进展声明。

原文摘要 · Abstract (English)

Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answering (VQA) benchmarks. We audit these claims and find them fundamentally compromised by data leakage at two hierarchical levels: patient-level leakage, where slides from the same case appear in both training and test folds, and institutional-level leakage, where different cases nonetheless share staining-batch and scanner signatures through a common Tissue Source Site (TSS). By tracing canonical slide, case, and TSS identifiers across major public resources, we document case level train test overlaps of 92.3~100% on TCGA-derived benchmarks, together with near-complete TSS overlap. We further demonstrate that both leakage levels are linearly decodable from foundation-model feature space, that they induce a measurable accuracy gap between leaked and audit-clean cases on a published checkpoint, and that across multiple published WSI VLMs, peak reported accuracies concentrate on the most heavily contaminated benchmarks. Therefore, the current WSI VQA evaluation cannot distinguish genuine multimodal reasoning from nearest-neighbor retrieval over memorized institutional and patient-specific artifacts. Finally, we outline concrete recommendations for contamination-free evaluation. By addressing benchmark construction, provenance disclosure, and automated overlap auditing, we aim to guide future research toward verifiable claims of progress.

医学影像数据泄露模型评估病理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。