用临床报告提炼推理链,提升肺部影像诊断模型的推理能力。
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
- 从真实报告中提取并优化推理链条,指导模型学习临床思维
- 在5.9万条数据上训练,推理能力比最强医疗模型高16%
- 开源完整数据集与评估工具,适合医学AI研究者使用
近期基于推理增强的大语言模型和多模态大语言模型在复杂任务中表现显著提升,但医疗AI模型常忽略临床实践中固有的结构化推理过程。本文提出ChestX-Reasoner,一种面向胸部影像诊断的多模态大语言模型,通过直接挖掘常规放射科报告中的推理链条,模拟放射科医生的逐步分析思路。我们构建了一个大规模数据集,从真实报告中提取并精炼推理路径。采用两阶段训练框架,结合监督微调与基于过程奖励的强化学习,使模型推理更贴近临床标准。提出RadRBench-CXR基准,包含59,000个视觉问答样本及301,000条临床验证的推理步骤,并设计RadRScore评估推理的真实性、完整性和有效性。ChestX-Reasoner在诊断准确率和推理能力上均优于现有医疗及通用领域多模态模型,相较最优医疗模型、最优通用模型及自身基线,在推理能力上分别提升16%、5.9%和18%,在结果准确率上分别提升3.3%、24%和27%。所有资源均已开源,推动医学推理多模态模型研究。
原文摘要 · Abstract (English)
Recent advances in reasoning-enhanced large language models (LLMs) and multimodal LLMs (MLLMs) have significantly improved performance in complex tasks, yet medical AI models often overlook the structured reasoning processes inherent in clinical practice. In this work, we present ChestX-Reasoner, a radiology diagnosis MLLM designed to leverage process supervision mined directly from clinical reports, reflecting the step-by-step reasoning followed by radiologists. We construct a large dataset by extracting and refining reasoning chains from routine radiology reports. Our two-stage training framework combines supervised fine-tuning and reinforcement learning guided by process rewards to better align model reasoning with clinical standards. We introduce RadRBench-CXR, a comprehensive benchmark featuring 59K visual question answering samples with 301K clinically validated reasoning steps, and propose RadRScore, a metric evaluating reasoning factuality, completeness, and effectiveness. ChestX-Reasoner outperforms existing medical and general-domain MLLMs in both diagnostic accuracy and reasoning ability, achieving 16%, 5.9%, and 18% improvements in reasoning ability compared to the best medical MLLM, the best general MLLM, and its base model, respectively, as well as 3.3%, 24%, and 27% improvements in outcome accuracy. All resources are open-sourced to facilitate further research in medical reasoning MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。