arXiv:2505.18087cs.CVcs.AI2025-05NeurIPS被引 5

构建胸部X光结构化诊断推理评测基准,揭示大模型真实医学推理能力。

CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays

  • 基于MIMIC-CXR-JPG数据集,自动提取解剖区域、测量指标等中间推理步骤
  • 包含18,988个问答对,1,200例病例,支持多阶段视觉定位与测量评估
  • 揭示当前最强12个大模型在结构化推理和泛化上仍严重不足

近年来,大型视觉语言模型(LVLMs)在医学任务中展现出潜力,如报告生成和视觉问答。然而,现有评测主要关注最终诊断答案,难以反映模型是否进行临床有意义的推理。为此,我们提出CheXStruct与CXReasonBench,一个基于公开MIMIC-CXR-JPG数据集的结构化推理流程与评测基准。CheXStruct可从胸部X光片中自动生成一系列中间推理步骤,包括解剖区域分割、解剖标志定位、诊断测量计算及临床阈值应用。CXReasonBench利用该流程,评估模型能否执行临床有效的推理步骤,并探究其在结构化引导下的学习能力,实现细粒度、透明化的诊断推理评估。基准涵盖12种诊断任务的18,988个问答对和1,200例病例,每例最多配4个视觉输入,支持多路径、多阶段评估,包括解剖区域选择和诊断测量的视觉定位。即使最强的12个被测LVLMs在结构化推理和泛化能力上仍表现不佳,常无法将抽象知识与解剖定位的视觉理解有效关联。代码已开源。

原文摘要 · Abstract (English)

Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks, such as report generation and visual question answering. However, existing benchmarks focus mainly on the final diagnostic answer, offering limited insight into whether models engage in clinically meaningful reasoning. To address this, we present CheXStruct and CXReasonBench, a structured pipeline and benchmark built on the publicly available MIMIC-CXR-JPG dataset. CheXStruct automatically derives a sequence of intermediate reasoning steps directly from chest X-rays, such as segmenting anatomical regions, deriving anatomical landmarks and diagnostic measurements, computing diagnostic indices, and applying clinical thresholds. CXReasonBench leverages this pipeline to evaluate whether models can perform clinically valid reasoning steps and to what extent they can learn from structured guidance, enabling fine-grained and transparent assessment of diagnostic reasoning. The benchmark comprises 18,988 QA pairs across 12 diagnostic tasks and 1,200 cases, each paired with up to 4 visual inputs, and supports multi-path, multi-stage evaluation including visual grounding via anatomical region selection and diagnostic measurements. Even the strongest of 12 evaluated LVLMs struggle with structured reasoning and generalization, often failing to link abstract knowledge with anatomically grounded visual interpretation. The code is available at https://github.com/ttumyche/CXReasonBench

医学影像结构化推理视觉语言模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。