病理视觉语言模型看似懂病理,实则可能靠文字套路答题。
Do Pathology Vision-Language Models Truly See Pathology?

- 设计新基准PathBind,检验模型是否真能结合图像与文本
- 多模型测试显示,准确率高但看图理解能力弱
- 适合研究医学AI可解释性与评估方法的学者
病理视觉语言模型(VLMs)近年发展迅速,常以病理VQA基准上的答对率评估性能。但我们发现当前评估存在三方面问题:1)视觉证据并非必需,例如Gemini-3-Pro在无图像输入下仍达53.5%平均准确率;2)领域训练提升准确率,但视觉-语义绑定未同步增强,相较Qwen2.5-VL-7B,Patho-R1-7B的多模态增益低5.8分,注意力交并比(IoU)低3.7分;3)实体级注意力分散且查询不敏感,在PathVG上不同实体查询的注意力图高度相关。这些问题导致对模型真实多模态能力的误判。为此,我们提出PathBind基准,包含2600个样本:PathBind-VQA(1500题,六维度)、PathBind-PTA(600题,来自私有病理教学图谱)、PathBind-Grounding(500个专家标注区域级样本),各部分经任务特化自动筛选与专家评审,减少文本捷径,强化实体-区域对应。在PathBind的VQA样本及五个现有病理VQA基准上评估18个代表性模型,并在PathBind-Grounding和PathVG上进一步评估10个模型。结果表明,当前病理VLM在答题表现与视觉-语义绑定之间仍存在显著差距。
原文摘要 · Abstract (English)
Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. However, we dig into current evaluations and identify three overlooked issues: 1) Visual evidence is not always necessary. For instance, Gemini-3-Pro achieves 53.5% average accuracy across 5 VQA benchmarks without any visual input. 2) Domain training can improve accuracy without proportional gains in visual binding. Compared with Qwen2.5-VL-7B, Patho-R1-7B exhibits a 5.8-point lower multimodal gain and a 3.7-point lower attention IoU. 3) Entity-level attention is diffuse and weakly query-specific. On PathVG, attention maps remain highly correlated across different entity queries. These issues can lead to substantial misjudgments of pathology VLMs' actual multimodal capabilities. To this end, we present PathBind, a benchmark comprising 2,600 samples: PathBind-VQA with 1,500 questions across six dimensions, PathBind-PTA with 600 questions from a private pathology teaching atlas, and PathBind-Grounding with 500 expert-curated region-level samples. Each component undergoes task-specific automated filtering and expert review to reduce textual shortcuts and improve entity-region correspondence. We evaluate 18 representative VLMs on VQA samples of PathBind and five existing pathology VQA benchmarks, and further evaluate 10 VLMs on PathBind-Grounding and PathVG. Results show that current pathology VLMs still exhibit a substantial gap between answer-side performance and visual-semantic binding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。