arXiv:2606.27264cs.CV2026-06

构建可追溯的3D肺部CT推理基准,让AI诊断像医生一样一步步说明依据。

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

论文配图:CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
图 1 · 摘自论文原文
  • 设计四阶段诊断流程还原放射科医生思考路径:理解任务→观察影像→推断诊断→生成答案。
  • 基于CT-RATE数据集构建76,177条经专家验证的结构化推理轨迹,覆盖开放/封闭问答与报告生成。
  • 联合临床医生设计评估框架,实现对每一步推理的可量化、可验证评价,助力可信医疗AI发展。

多模态大模型在医学影像中的推理展现出巨大潜力,但其推理过程常为自由文本,仅凭最终答案评判,难以解释和验证,尤其在三维放射学中,诊断需能追溯到扫描证据。现有胸部CT问答数据集将专家报告简化为仅含答案的配对,丢失了从发现到结论的推理链条,并省略了临床医生依赖的患者病史。因此,具备结构化推理能力的3D胸部CT多模态大模型仍难实现,既缺乏训练所需的结构化监督,也无验证推理可信性的标准流程。本文提出CORTEX(Clinically Organized Reasoning and sTructured EXplanation),一个面向3D胸部CT的结构化推理基准。针对每个问题,CORTEX恢复缺失的推理过程,构建符合放射科医生工作流的四阶段诊断轨迹:任务理解、视觉观察、诊断推理、答案合成。利用具备广泛医学与通用知识的前沿大模型生成这些轨迹,并通过结合自动评分与专家放射科医生评审的阶段级评估协议进行过滤与验证。关键在于,推理结构与评估标准均与临床医生密切协作设计。基于公开可用的无推理标注的胸部CT数据集CT-RATE,CORTEX包含76,177条经验证的推理轨迹,涵盖开放式视觉问答、封闭式视觉问答及报告生成任务,为构建和评估可信3D胸部CT多模态大模型提供了必需的结构化监督与阶段级评估协议。数据集与评估代码将在论文接收后公开。

原文摘要 · Abstract (English)

Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer, making it hard to interpret and verify, especially in 3D radiology, where a diagnosis should be traceable to evidence in the scan. Existing chest CT question-answering datasets compound this by reducing expert radiology reports to answer-only pairs, dropping the reasoning that links findings to conclusions and omitting the patient history clinicians rely on. As a result, reasoning-capable 3D chest CT MLLMs remain out of reach, as neither the structured supervision needed to train them nor the protocol needed to verify their reasoning yet exists. We introduce CORTEX (Clinically Organized Reasoning and sTructured EXplanation), a structured reasoning benchmark for 3D chest CT. For each question, CORTEX restores the missing reasoning as a four-stage diagnostic trace mirroring a radiologist's workflow: task understanding, visual observation, diagnostic reasoning, and answer synthesis. We generate these traces using frontier large language models with broad medical and general-domain knowledge, then filter and verify them with a stage-level evaluation protocol combining automated rubric scoring with expert radiologist review. Crucially, both the reasoning structure and evaluation rubrics are designed in close collaboration with clinicians. Built on CT-RATE, a large, publicly available chest CT dataset without reasoning annotations, CORTEX comprises 76,177 validated reasoning traces across open-ended VQA, closed-ended VQA, and report generation, providing both the structured supervision and the stage-level evaluation protocol needed to build and evaluate trustworthy reasoning models for 3D chest CT. Our dataset and evaluation code will be made publicly available upon acceptance.

医疗AI推理链3D影像可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。