arXiv:2608.26382cs.CV2026-08

首个兽医病理视觉语言模型评测基准,填补非人类病理空白

VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology

论文配图:VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology
图 1 · 摘自论文原文
  • 由兽医病理专家手工标注1251个问题,覆盖419张大鼠组织切片
  • 发现前沿模型在兽医病理上误诊正常组织风险高,领域专训仍关键
  • 适合关注医疗AI泛化性、兽医病理或药物安全评估的研究者

病理视觉语言模型发展迅速,但现有基准多聚焦人类组织,尤其是肿瘤病理,忽略了非人类病理。这一缺口在毒理病理中尤为显著,因其依赖实验动物组织显微检查进行新药安全性评估。为此,我们提出VIPER,首个由专家标注的兽医病理视觉语言模型评测基准。该数据集包含419张H&E染色大鼠组织切片,涵盖七个器官系统,对应1251个问题,形式包括多项选择、KPrim和自由文本。所有问题均由持证兽医病理学家设计并验证。我们共评测了16个模型,包括两个新提出的兽医病理专用模型、七个专注人类病理的模型以及七个通用前沿模型。结果揭示了兽医与人类病理间存在显著领域差距,暴露了前沿模型对正常组织的过度诊断风险,并证实领域特定训练对视觉-语言预测至关重要。VIPER数据及评估代码已开源:https://github.com/mahmoodlab/viper。

原文摘要 · Abstract (English)

Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncology, leaving non-human pathology largely unaddressed. This gap is especially important in toxicologic pathology, where microscopic tissue examination of laboratory animals is a core component of preclinical drug safety assessment. To address it, we introduce VIPER, the first expert-curated benchmark for vision-language model evaluation in toxicologic pathology. VIPER contains 1,251 questions associated with 419 H&E-stained rat histology images across seven organ systems, covering multiple-choice, KPrim, and free-text formats. All questions were curated and validated by board-certified veterinary pathologists. In total, we benchmarked 16 models, including two newly introduced veterinary-pathology models, seven human pathology-specialized models, and seven general-purpose frontier models. The results identify a substantial domain gap between veterinary and human pathology, expose the risk of over-diagnosis of normal tissue in frontier models, and show that domain-specific training remains critical for visually grounded predictions. VIPER data and evaluation code are available at https://github.com/mahmoodlab/viper.

兽医病理视觉语言基准测试AI医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。