让医疗异常检测模型推理更连贯可信,提升临床AI可解释性。
MedAD-R1: Eliciting Consistent Reasoning in Interpretible Medical Anomaly Detection via Consistency-Reinforced Policy Optimization
- 用分步训练框架注入医学知识并强化推理与答案的一致性
- 在38K规模数据集上超越基线10%以上,推理逻辑更连贯
- 适合追求可解释医疗AI的临床研究者和开发者
医学异常检测(MedAD)利用大模型分析医疗影像并回答问题,有望提升诊断准确率。但现有方法依赖简单碎片化数据进行监督微调,限制了模型的合理推理与多模态泛化能力。为此,我们构建了首个大规模、多中心、多模态的医学异常检测基准MedAD-38K,包含诊断链式思维(CoT)标注与结构化视觉问答对。在此基础上,提出两阶段训练框架:第一阶段认知注入(Cognitive Injection)通过监督微调注入基础医学知识,并引导模型采用‘先思考后作答’范式;第二阶段引入一致性组相对策略优化(Con-GRPO),设计关键一致性奖励机制,确保推理过程与最终诊断逻辑一致。所提出的MedAD-R1模型在MedAD-38K上达到当前最优性能,相比强基线提升超过10%。其优势在于生成透明且逻辑自洽的推理路径,为临床决策支持中的AI可信度与可解释性提供新思路。
原文摘要 · Abstract (English)
Medical Anomaly Detection (MedAD) presents a significant opportunity to enhance diagnostic accuracy using Large Multimodal Models (LMMs) to interpret and answer questions based on medical images. However, the reliance on Supervised Fine-Tuning (SFT) on simplistic and fragmented datasets has hindered the development of models capable of plausible reasoning and robust multimodal generalization. To overcome this, we introduce MedAD-38K, the first large-scale, multi-modal, and multi-center benchmark for MedAD featuring diagnostic Chain-of-Thought (CoT) annotations alongside structured Visual Question-Answering (VQA) pairs. On this foundation, we propose a two-stage training framework. The first stage, Cognitive Injection, uses SFT to instill foundational medical knowledge and align the model with a structured think-then-answer paradigm. Given that standard policy optimization can produce reasoning that is disconnected from the final answer, the second stage incorporates Consistency Group Relative Policy Optimization (Con-GRPO). This novel algorithm incorporates a crucial consistency reward to ensure the generated reasoning process is relevant and logically coherent with the final diagnosis. Our proposed model, MedAD-R1, achieves state-of-the-art (SOTA) performance on the MedAD-38K benchmark, outperforming strong baselines by more than 10\%. This superior performance stems from its ability to generate transparent and logically consistent reasoning pathways, offering a promising approach to enhancing the trustworthiness and interpretability of AI for clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。