提升3D医学影像模型的临床推理能力,让诊断更准确透明。
Med3D-R1: Incentivizing Clinical Reasoning in 3D Medical Vision-Language Models for Abnormality Diagnosis
- 分两阶段训练:先对齐影像与文本特征,再用强化学习优化推理过程。
- 在两个3D诊断数据集上达到41.92%和44.99%的最高准确率。
- 特别关注异常区域识别,适合医疗AI研发与临床辅助系统应用。
由于体积分割医学影像的固有复杂性、模型对报告表面模式的过拟合倾向以及缺乏可解释性感知的奖励设计,构建具备稳健临床推理能力的3D视觉语言模型仍面临挑战。本文提出Med3D-R1,一种基于强化学习的两阶段训练框架:监督微调(SFT)与强化学习(RL)。在SFT阶段,引入残差对齐机制以弥合高维3D特征与文本嵌入之间的差距,并采用异常重加权策略突出临床相关信息词元,降低报告中的结构偏差。在RL阶段,重新设计一致性奖励,显式促进连贯、逐步的诊断推理。我们在两个3D诊断基准数据集CT-RATE和RAD-ChestCT上评估该方法,其在多选视觉问答任务中分别取得41.92%和44.99%的最新最优准确率。结果表明,模型在异常检测与临床推理能力上均有提升,优于现有方法。整体而言,该方法有望通过增强3D医学视觉语言系统的可靠性与透明性,改善真实诊疗流程。
原文摘要 · Abstract (English)
Developing 3D vision-language models with robust clinical reasoning remains a challenge due to the inherent complexity of volumetric medical imaging, the tendency of models to overfit superficial report patterns, and the lack of interpretability-aware reward designs. In this paper, we propose Med3D-R1, a reinforcement learning framework with a two-stage training process: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). During SFT stage, we introduce a residual alignment mechanism to bridge the gap between high-dimensional 3D features and textual embeddings, and an abnormality re-weighting strategy to emphasize clinically informative tokens and reduce structural bias in reports. In RL stage, we redesign the consistency reward to explicitly promote coherent, step-by-step diagnostic reasoning. We evaluate our method on medical multiple-choice visual question answering using two 3D diagnostic benchmarks, CT-RATE and RAD-ChestCT, where our model attains state-of-the-art accuracies of 41.92\% on CT-RATE and 44.99\% on RAD-ChestCT. These results indicate improved abnormality diagnosis and clinical reasoning and outperform prior methods on both benchmarks. Overall, our approach holds promise for enhancing real-world diagnostic workflows by enabling more reliable and transparent 3D medical vision-language systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。