构建病理图像多尺度理解基准,揭示大模型在细粒度视觉分析中的短板。
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

- 基于23个公开数据集设计双视野评估任务,覆盖局部与整体视图。
- 18款模型在细粒度识别、数量估计等任务上表现普遍不足。
- 适合开发医疗视觉大模型的研究者用于系统性评估多尺度理解能力。
多模态大语言模型(MLLMs)在病理图像分析中应用日益广泛,但现有病理学主流基准主要评估最终诊断结果、描述或报告,难以揭示模型是否真正理解病灶分析所需的多尺度视觉内容。为此,我们提出PathVU——一个以视觉为中心的细粒度、多尺度视觉理解基准,涵盖23个公开病理影像数据集,包含人工标注与空间标注。该基准在两种视野下评估:区域视野(Region FOV)用于高分辨率局部区域,全片视野(Slide FOV)用于宏观整体切片。通过将原始标注转化为确定性任务目标,实现对区域定位、视觉识别、数量估计、空间推理及上下文不足判断的程序化评分。基准包含14项VQA风格任务,共61,673张图像、308,070个样本,覆盖28个器官与7,253,526个标注。对18种代表性通用型、医学领域及病理专用MLLM的评估显示,即使先进模型在多尺度病理图像的细粒度视觉任务上仍存在显著局限。PathVU为开发与评估具备明确多尺度视觉理解能力的病理学MLLM提供了可复现的基础。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathVU, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from 23 public pathology imaging datasets with human-supervised labels and spatial annotations, PathVU evaluates MLLM understanding in two fields of view: Region FOV for high-resolution local regions and Slide FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathVU enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains 14 VQA-style tasks, 61,673 images, and 308,070 samples across 28 organs and 7,253,526 annotations. Evaluating 18 representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathVU provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。