首个多中心病理AI助手基准,测试大模型辅助诊断能力。
DALPHIN: Benchmarking Digital Pathology AI Copilots Against Pathologists on an Open Multicentric Dataset

- 构建多中心开放数据集,覆盖130种疾病与14个亚专科。
- 病理专用模型PathChat在4项任务中媲美专家水平。
- 适合医学AI研究者、病理医生及评测平台开发者使用。
具备视觉问答能力的通用基础模型正兴起于数字病理领域。为评估其在常规诊断中的辅助潜力,我们创建了首个多中心开放基准DALPHIN,包含来自300例、6个国家、14个亚专科的1236张图像,涵盖130种从罕见到常见的诊断。同时引入31位来自10个国家、经验各异的病理医生作为人类性能基准。评估了两个通用模型(GPT-5、Gemini 2.5 Pro)和一个病理专用模型(PathChat+)在序列与独立回答生成下的表现。结果显示,PathChat在六项任务中有四项与专家水平无显著差异,Gemini在两项中表现相当,GPT仅一项达标。DALPHIN已公开发布,真实标签受控访问,以支持长期可靠评测。数据、方法与评估平台可通过 dalphin.grand-challenge.org 获取。
原文摘要 · Abstract (English)
Foundation models with visual question answering capabilities for digital pathology are emerging. Such unprecedented technology requires independent benchmarking to assess its potential in assisting pathologists in routine diagnostics. We created DALPHIN, the first multicentric open benchmark for pathology AI copilots, comprising 1236 images from 300 cases, spanning 130 rare to common diagnoses, 6 countries, and 14 subspecialties. The DALPHIN design and dataset are introduced alongside a human performance benchmark of 31 pathologists from 10 countries with varying expertise. We report results for two general-purpose (GPT-5, Gemini 2.5 Pro) and one pathology-specific copilot (PathChat+) for sequential and independent answer generation. We observed no statistically significant difference from expert-level performance in four of six tasks for PathChat, 2/6 tasks for Gemini, and 1/6 tasks for GPT. DALPHIN is publicly released with sequestered, indirectly accessible ground truth to foster robust and enduring benchmarking. Data, methods, and the evaluation platform are accessible through dalphin.grand-challenge.org.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。