首个眼科多模态渐进推理基准,评测模型临床决策能力
X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

- 构建六阶段渐进推理链,覆盖从图像质量到诊断的完整流程
- 整合六种眼科影像模态,评估跨模态信息融合能力
- 涵盖52种眼病,26,415张图像,适合评估医疗大模型临床推理
尽管多模态大语言模型(MLLMs)取得显著进展,其在多模态诊断中的临床推理能力仍缺乏系统评估。现有基准多为单模态数据,无法检验临床实践中至关重要的渐进式推理与跨模态整合能力。本文提出首个面向眼科诊断的跨模态渐进临床推理基准X-PCR,通过完整的诊疗流程评估MLLMs,包含两个任务:1)涵盖图像质量评估至临床决策的六阶段渐进推理链;2)整合六种成像模态的跨模态推理任务。基准包含26,415张图像和177,868对专家验证的VQA问答对,来自51个公开数据集,覆盖52种眼科疾病。对21个MLLMs的评估揭示了其在渐进推理与跨模态整合上的显著缺陷。数据集与代码已开源:https://github.com/CVI-SZU/X-PCR。
原文摘要 · Abstract (English)
Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely unexamined. Current benchmarks, mostly single-modality data, can't evaluate progressive reasoning and cross-modal integration essential for clinical practice. We introduce the Cross-Modality Progressive Clinical Reasoning (X-PCR) benchmark, the first comprehensive evaluation of MLLMs through a complete ophthalmology diagnostic workflow, with two reasoning tasks: 1) a six-stage progressive reasoning chain spanning image quality assessment to clinical decision-making, and 2) a cross-modality reasoning task integrating six imaging modalities. The benchmark comprises 26,415 images and 177,868 expert-verified VQA pairs curated from 51 public datasets, covering 52 ophthalmic diseases. Evaluation of 21 MLLMs reveals critical gaps in progressive reasoning and cross-modal integration. Dataset and code: https://github.com/CVI-SZU/X-PCR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。