arXiv:2607.19261cs.CVcs.AI2026-07被引 3

评测视觉语言模型在全切片病理图像中获取与推理证据的能力

EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology

论文配图:EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology
图 1 · 摘自论文原文
  • 构建诊断树结构,评估多尺度证据获取与推理能力
  • 19个模型在多尺度推理上准确率超93%,但定位能力不足
  • 揭示当前模型依赖预提取证据,自主探索能力弱

全切片图像(WSI)诊断需识别关键区域、跨分辨率分析并整合多尺度证据。然而现有病理基准大多基于预裁剪图像或预提取特征,难以评估模型从千兆像素级WSI中自主获取证据的能力。本文提出EviPathBench,一个用于评估视觉语言模型(VLMs)在全切片病理中证据获取与推理的基准。该基准涵盖四项能力:图像到文本匹配(证据解释)、文本到图像检索(证据验证)、诊断区域定位(证据获取)和多尺度推理(证据整合)。基准采用诊断树结构,连接不同放大倍数下的嵌套区域与对应的尺度特异性发现及病理诊断。包含1,822个TCGA WSI和17,135条由十位认证病理科医生标注的诊断路径;另设190个乳腺癌病例的私有队列,用于评估自主全切片探索能力。评估了19个覆盖通用、医学及病理专用类别的VLMs,以及一个纯文本参考模型。领先开源模型在多尺度推理上准确率超过93%,在跨模态匹配任务中表现均超50%。但诊断区域定位仍具挑战:最优文本引导的平均交并比低于0.09,劣于中心启发式方法。自主探索中,无条件命中率从低倍镜的0.522降至中倍镜的0.185和高倍镜的0.020。结果揭示了模型在已有证据上的推理能力与从原始图像中自主获取证据之间存在显著差距。EviPathBench为统一衡量和提升这两方面能力提供了框架。

原文摘要 · Abstract (English)

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most pathology benchmarks evaluate models on pre-cropped patches or pre-extracted slide features, leaving their ability to acquire evidence from gigapixel WSIs largely untested. We introduce EviPathBench, a benchmark for evaluating evidence acquisition and reasoning in vision-language models (VLMs) for whole-slide pathology. It evaluates four capabilities: image-to-text matching for evidence interpretation, text-to-image retrieval for evidence verification, diagnostic-region localization for evidence acquisition, and multi-scale reasoning for evidence integration. The benchmark is organized as a diagnostic tree linking nested regions across magnifications with scale-specific findings and path-level diagnoses. It contains 1,822 TCGA WSIs and 17,135 diagnostic paths annotated by ten board-certified pathologists. A private cohort of 190 breast cancer WSIs with detailed annotations further evaluates autonomous whole-slide exploration. We evaluate 19 VLMs spanning general-purpose, medical, and pathology-specialized families, plus one text-only reference model. Leading open-weight models achieve over 93% accuracy in multi-scale reasoning and over 50% in both cross-modal matching tasks. In contrast, diagnostic-region localization remains challenging: the best text-guided mean intersection-over-union is below 0.09, underperforming a center-based heuristic. During autonomous exploration, the unconditional hit rate drops from 0.522 at low magnification to 0.185 at intermediate magnification and 0.020 at high magnification. These results reveal a pronounced gap between reasoning over curated evidence and acquiring it from WSIs. EviPathBench provides a unified framework for measuring and improving both capabilities.

病理图像多尺度推理视觉语言模型证据获取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。