arXiv:2512.00818cs.AIcs.CV2025-12被引 13

构建医疗多模态复杂推理的细粒度评测基准,测试模型真实临床能力。

Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

  • 将医学多模态推理拆解为视觉理解与多步推理,实现精准评估
  • 覆盖11个器官系统、12种影像模态,含20653组高质量VQA数据
  • 揭示通用模型在罕见病例上表现优于专用模型,暴露长尾泛化缺陷

多模态大模型正进入临床流程,但其复杂医学推理能力尚不明确。我们提出Med-CMR,一个细粒度的医学复杂多模态推理评测基准。该基准具有三大核心特性:1)系统性能力分解,将医学多模态推理拆分为细粒度的视觉理解与多步推理,支持靶向评估;2)高挑战性任务设计,视觉理解涵盖小物体检测、细粒度判别与空间理解三个维度,推理覆盖时间预测、因果推理、长尾泛化与多源融合四种临床场景;3)广泛且高质量的数据覆盖,包含20,653组视觉问答(VQA)对,覆盖11个器官系统与12种影像模态,经两阶段(专家+模型辅助)审核确保临床真实性。我们评估了18个先进MLLMs,结果显示GPT-5在多项选择题(MCQ)上达到57.81准确率,开放问答得分为48.70,优于Gemini 2.5 Pro(49.87 MCQ,45.98开放得分)与开源模型Qwen3-VL-235B-A22B(49.34 MCQ,42.62开放得分)。然而,专业医学模型并未稳定超越强通用模型,长尾泛化成为主要失效模式。Med-CMR因此为视觉-推理融合与罕见病例鲁棒性提供了压力测试,并为未来临床系统提供严谨评估标准。

原文摘要 · Abstract (English)

MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from existing counterparts by three core features: 1) Systematic capability decomposition, splitting medical multimodal reasoning into fine-grained visual understanding and multi-step reasoning to enable targeted evaluation; 2) Challenging task design, with visual understanding across three key dimensions (small-object detection, fine-detail discrimination, spatial understanding) and reasoning covering four clinically relevant scenarios (temporal prediction, causal reasoning, long-tail generalization, multi-source integration); 3) Broad, high-quality data coverage, comprising 20,653 Visual Question Answering (VQA) pairs spanning 11 organ systems and 12 imaging modalities, validated via a rigorous two-stage (human expert + model-assisted) review to ensure clinical authenticity. We evaluate 18 state-of-the-art MLLMs with Med-CMR, revealing GPT-5 as the top-performing commercial model: 57.81 accuracy on multiple-choice questions (MCQs) and a 48.70 open-ended score, outperforming Gemini 2.5 Pro (49.87 MCQ accuracy, 45.98 open-ended score) and leading open-source model Qwen3-VL-235B-A22B (49.34 MCQ accuracy, 42.62 open-ended score). However, specialized medical MLLMs do not reliably outperform strong general models, and long-tail generalization emerges as the dominant failure mode. Med-CMR thus provides a stress test for visual-reasoning integration and rare-case robustness in medical MLLMs, and a rigorous yardstick for future clinical systems.

医疗AI多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。