首个面向罕见病的多模态多图像医学评测基准,填补临床决策空白。
MMRareBench: A Rare-Disease Multimodal and Multi-Image Medical Benchmark

- 构建罕见病多模态多图像评测框架,涵盖诊断、治疗规划等四类任务。
- 包含1756个问答对与7958张医学图像,基于Orphanet知识体系标注。
- 揭示医疗模型在罕见病多图整合中能力下降,提示微调可能削弱复合推理能力。
多模态大语言模型在常见病临床任务中表现优异,但在罕见病场景下的性能仍缺乏系统评估。罕见病诊疗中,医生常缺乏既往经验,必须依赖病例级证据作出判断。现有基准大多聚焦常见病、单图场景,未能覆盖多模态与多图像证据融合在数据稀缺下的评估需求。本文提出MMRareBench,据我们所知首个联合评估罕见病多模态与多图像临床能力的基准,包含四个与临床流程对齐的任务:诊断、治疗规划、跨图像证据对齐和检查建议。该基准共包含1756个问答对及7958张来自PMC病例报告的医学图像,采用Orphanet锚定的本体对齐、任务特异性泄露控制、基于证据的标注及两级评估协议。对23个MLLM的系统评估显示,其能力呈现碎片化特征,治疗规划性能普遍偏低;尽管医疗领域模型在诊断任务上表现良好,但在多图像任务中显著落后于通用模型。这一现象符合‘能力稀释效应’:医疗微调虽可缩小诊断差距,但可能削弱罕见病证据整合所需的组合式多图像推理能力。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely untested. In rare-disease scenarios, clinicians often lack prior clinical knowledge, forcing them to rely strictly on case-level evidence for clinical judgments. Existing benchmarks predominantly evaluate common-condition, single-image settings, leaving multimodal and multi-image evidence integration under rare-disease data scarcity systematically unevaluated. We introduce MMRareBench, to our knowledge the first rare-disease benchmark jointly evaluating multimodal and multi-image clinical capability across four workflow-aligned tracks: diagnosis, treatment planning, cross-image evidence alignment, and examination suggestion. The benchmark comprises 1,756 question-answer pairs with 7,958 associated medical images curated from PMC case reports, with Orphanet-anchored ontology alignment, track-specific leakage control, evidence-grounded annotations, and a two-level evaluation protocol. A systematic evaluation of 23 MLLMs reveals fragmented capability profiles and universally low treatment-planning performance, with medical-domain models trailing general-purpose MLLMs substantially on multi-image tracks despite competitive diagnostic scores. These patterns are consistent with a capacity dilution effect: medical fine-tuning can narrow the diagnostic gap but may erode the compositional multi-image capability that rare-disease evidence integration demands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。