arXiv:2601.08758eess.IVcs.CV2026-01中稿 · ICLR被引 3

首个评估医学影像中多模态大模型推理链的基准测试

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

  • 构建涵盖24种检查类型的多层级难度数据集
  • 提出正确性、效率、影响与一致性四项专属评估指标
  • 揭示当前医学大模型推理不可靠,适合医疗AI研究者使用

链式思维(CoT)推理通过引导逐步中间推理,在提升大语言模型性能方面表现优异,近期已扩展至多模态大语言模型(MLLMs)。在医学领域,诊断依赖于细微的视觉线索与序列化推理,CoT与临床思维天然契合。然而,现有医学图像理解基准普遍关注最终答案,忽视推理路径。这种黑箱推理缺乏可靠判断依据,难以辅助医生诊断。为此,我们提出M3CoTBench基准,专门评估医学图像理解中CoT推理的正确性、效率、影响与一致性。该基准包含:1)覆盖24种检查类型的多样化、多层级难度数据集;2)13个不同难度的任务;3)针对临床推理设计的四类CoT专用评估指标;4)对多个MLLMs的性能分析。M3CoTBench系统性评估了多种医学影像任务中的CoT推理,揭示了当前MLLMs在生成可靠、临床可解释推理方面的局限性,旨在推动透明、可信且诊断准确的医疗AI系统发展。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step-by-step intermediate reasoning, and recent advances have extended this paradigm to Multimodal Large Language Models (MLLMs). In the medical domain, where diagnostic decisions depend on nuanced visual cues and sequential reasoning, CoT aligns naturally with clinical thinking processes. However, current benchmarks for medical image understanding generally focus on the final answer while ignoring the reasoning path. Such opaque reasoning processes lack reliable bases for judgment, making it difficult to assist doctors in diagnosis. To address this gap, we introduce a new M3CoTBench benchmark specifically designed to evaluate the correctness, efficiency, impact, and consistency of CoT reasoning in medical image understanding. M3CoTBench features 1) a diverse, multi-level difficulty dataset covering 24 examination types, 2) 13 varying-difficulty tasks, 3) a suite of CoT-specific evaluation metrics (correctness, efficiency, impact, and consistency) tailored to clinical reasoning, and 4) a performance analysis of multiple MLLMs. M3CoTBench systematically evaluates CoT reasoning across diverse medical imaging tasks, revealing current limitations of MLLMs in generating reliable and clinically interpretable reasoning, and aims to foster the development of transparent, trustworthy, and diagnostically accurate AI systems for healthcare. Project page at https://juntaojianggavin.github.io/projects/M3CoTBench/.

医学影像推理链多模态模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。