首个开放的组织病理学多模态模型综合评测基准,揭示现有模型脆弱性。
How Good is my Histopathology Vision-Language Foundation Model? A Holistic Benchmark
- 构建覆盖26个器官、31种癌症的多源图像-文本数据集
- 发现多数模型对文字描述变化敏感,准确率下降最高达25%
- 揭示模型校准差、抗干扰能力弱,不适合临床部署
近年来,组织病理学视觉-语言基础模型(VLMs)因在多种下游任务中表现优异且具备强泛化能力而受到关注。然而,现有组织病理学评测基准大多为单模态,或在临床任务、器官种类、设备来源多样性方面受限,且由于患者隐私问题难以公开。因此,缺乏一个统一、全面的评估体系来反映真实临床场景。为此,我们提出HistoVL,一个完全开源的综合性评测基准,包含使用多达11种不同设备采集的图像,配以结合病种名称和多样病理描述的定制化文本标注。该数据集涵盖26个器官、31种癌症类型,来自14个异质患者队列,共超过500万张切片图像,源自41,000余张全切片图像(WSIs),覆盖多种放大倍数。我们系统评估了现有组织病理学VLMs在HistoVL上的表现,模拟专家在真实临床中的多样化任务。分析发现:多数现有模型对文本变化极为敏感,如转移灶检测任务中平衡准确率最高下降25%;对对抗攻击鲁棒性差;模型校准不佳,表现为高ECE值与低预测置信度,均可能影响其临床应用可靠性。
原文摘要 · Abstract (English)
Recently, histopathology vision-language foundation models (VLMs) have gained popularity due to their enhanced performance and generalizability across different downstream tasks. However, most existing histopathology benchmarks are either unimodal or limited in terms of diversity of clinical tasks, organs, and acquisition instruments, as well as their partial availability to the public due to patient data privacy. As a consequence, there is a lack of comprehensive evaluation of existing histopathology VLMs on a unified benchmark setting that better reflects a wide range of clinical scenarios. To address this gap, we introduce HistoVL, a fully open-source comprehensive benchmark comprising images acquired using up to 11 various acquisition tools that are paired with specifically crafted captions by incorporating class names and diverse pathology descriptions. Our Histo-VL includes 26 organs, 31 cancer types, and a wide variety of tissue obtained from 14 heterogeneous patient cohorts, totaling more than 5 million patches obtained from over 41K WSIs viewed under various magnification levels. We systematically evaluate existing histopathology VLMs on Histo-VL to simulate diverse tasks performed by experts in real-world clinical scenarios. Our analysis reveals interesting findings, including large sensitivity of most existing histopathology VLMs to textual changes with a drop in balanced accuracy of up to 25% in tasks such as Metastasis detection, low robustness to adversarial attacks, as well as improper calibration of models evident through high ECE values and low model prediction confidence, all of which can affect their clinical implementation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。