构建临床多模态评估基准,区分模型推理与文献检索能力
CURE: A Multimodal Benchmark for Clinical Understanding and Retrieval Evaluation
- 设计500个临床案例,配对医生引用的参考文献,分离评估推理与检索
- 模型有参考文献时诊断准确率达73.4%,独立检索时降至25.4%
- 适合医疗AI研发者、评测人员及多模态模型优化研究者使用
多模态大语言模型(MLLMs)在临床诊断中展现巨大潜力,该领域需整合复杂视觉与文本信息,并参考权威医学文献。然而现有基准多聚焦端到端问答,难以区分模型的多模态推理能力与证据检索能力。我们提出临床理解与检索评估基准CURE,包含500个映射至医师引用文献的多模态临床案例,可在受控证据环境下评估推理与检索表现。我们在封闭式与开放式诊断任务中评估主流MLLMs在不同证据获取模式下的表现。结果显示显著差异:当提供医师参考文献时,先进模型在鉴别诊断中最高可达73.4%准确率;而依赖自主检索时,准确率最低仅25.4%。这一差距凸显了有效整合多模态临床证据与精准检索支持文献的双重挑战。CURE已公开于https://github.com/yanniangu/CURE。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) demonstrate considerable potential in clinical diagnostics, a domain that inherently requires synthesizing complex visual and textual data alongside consulting authoritative medical literature. However, existing benchmarks primarily evaluate MLLMs in end-to-end answering scenarios. This limits the ability to disentangle a model's foundational multimodal reasoning from its proficiency in evidence retrieval and application. We introduce the Clinical Understanding and Retrieval Evaluation (CURE) benchmark. Comprising $500$ multimodal clinical cases mapped to physician-cited reference literature, CURE evaluates reasoning and retrieval under controlled evidence settings to disentangle their respective contributions. We evaluate state-of-the-art MLLMs across distinct evidence-gathering paradigms in both closed-ended and open-ended diagnosis tasks. Evaluations reveal a stark dichotomy: while advanced models demonstrate clinical reasoning proficiency when supplied with physician reference evidence (achieving up to $73.4\%$ accuracy on differential diagnosis), their performance substantially declines (as low as $25.4\%$) when reliant on independent retrieval mechanisms. This disparity highlights the dual challenges of effectively integrating multimodal clinical evidence and retrieving precise supporting literature. CURE is publicly available at https://github.com/yanniangu/CURE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。