构建医疗影像多维评估框架,揭示大模型推理可靠性问题。
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging

- 设计两级构建流程,实现细粒度多维度评估
- 33个模型评测中,灵枢-32B表现最佳
- 发现模型捷径行为与诊断性能正相关,警示临床可信风险
多模态大语言模型在医学影像领域的应用潜力催生了与真实临床实践对齐的系统性、严谨评估框架需求。现有方法仅报告单一或粗粒度指标,缺乏专业化临床支持所需的细致评估能力,且无法检验推理机制的可靠性。为此,我们提出向多维、细粒度、深入评估范式转变。基于为该范式设计的两阶段系统构建流程,我们构建了MedRCube。我们对33个MLLM进行了基准测试,其中*Lingshu-32B*表现达到顶尖水平。关键的是,MedRCube揭示了以往评估方式无法捕捉的一系列显著洞见。此外,我们引入可信度评估子集以量化推理可信度,发现捷径行为与诊断任务性能存在高度显著的正相关,对临床可信赖部署提出警示。相关资源见https://github.com/F1mc/MedRCube。
原文摘要 · Abstract (English)
The potential of Multimodal Large Language Models (MLLMs) in domain of medical imaging raise the demands of systematic and rigorous evaluation frameworks that are aligned with the real-world medical imaging practice. Existing practices that report single or coarse-grained metrics are lack the granularity required for specialized clinical support and fail to assess the reliability of reasoning mechanisms. To address this, we propose a paradigm shift toward multidimensional, fine-grained and in-depth evaluation. Based on a two-stage systematic construction pipeline designed for this paradigm, we instantiate it with MedRCube. We benchmark 33 MLLMs, \textit{Lingshu-32B} achieve top-tier performance. Crucially, MedRCube exposes a series of pronounced insights inaccessible under prior evaluation settings. Furthermore, we introduce a credibility evaluation subset to quantify reasoning credibility, uncover a highly significant positive association between shortcut behavior and diagnostic task performance, raising concerns for clinically trustworthy deployment. The resources of this work can be found at https://github.com/F1mc/MedRCube.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。