让医学AI诊断像医生一样说清楚理由,还能通过临床验证。
Toward Clinically Explainable AI for Medical Diagnosis: A Foundation Model with Human-Compatible Reasoning via Reinforcement Learning
- 用强化学习训练模型,让它一步步说出诊断依据
- 在胸部X光任务中,生成报告和问答表现优于现有模型
- 临床专家更认可它的推理过程,适合需要解释的医疗场景
医学AI在临床应用中受限于其黑箱特性,难以让医生验证决策依据。为此,我们提出DeepMedix-R1,一个用于胸部X光(CXR)解读的基础模型(FM),不仅能准确诊断,还能基于具体视觉证据生成透明、分步的推理过程。方法上采用序列化训练策略:先指令微调,再通过冷启动激发推理能力,最后使用基于视觉证据的强化学习精细优化模型,使其诊断结果与推理路径均符合临床合理性。定量评估显示,DeepMedix-R1在报告生成和视觉问答任务中显著优于先进基础模型。我们还构建了Report Arena——一个基于大语言模型的基准,结果显示DeepMedix-R1在输出质量上排名第一。最关键是,临床专家正式评审表明,其生成的推理过程明显优于广泛使用的Qwen2.5-VL-7B模型,证实其更强的可解释性与临床实用性。
原文摘要 · Abstract (English)
The clinical adoption of artificial intelligence (AI) in medical diagnostics is critically hampered by its black-box nature, which prevents clinicians from verifying the rationale behind automated decisions. To overcome this fundamental barrier, we introduce DeepMedix-R1, a foundation model (FM) for chest X-ray (CXR) interpretation that generates not only accurate diagnoses but also a transparent, step-by-step reasoning process grounded in specific visual evidence. Our methodology employs a sequential training strategy, beginning with instruction fine-tuning, followed by a cold-start phase to elicit reasoning capabilities. Critically, we then implement reinforcement learning with grounded rewards to meticulously refine the model, aligning both its diagnostic outputs and its reasoning pathways with clinical plausibility. Quantitative assessments show that DeepMedix-R1 substantially outperforms advanced FMs, achieving improvements in report generation and visual question answering tasks. We also introduce Report Arena, a novel LLM-based benchmark that ranks DeepMedix-R1 first among competing models for output quality. Most significantly, a formal review by clinical experts reveals a profound preference for DeepMedix-R1's generated reasoning over the broadly adopted Qwen2.5-VL-7B model, confirming its superior interpretability and clinical utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。