用强化学习让AI理解医生隐含问诊并精准定位病灶。
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
- 分离推理与分割,用强化学习优化临床推理过程
- 在14,000样本数据集上实现像素级定位,准确率领先
- 适合医疗AI研发者、影像科医生及可解释性研究者
在医学影像中精确定位病灶对诊断与治疗规划至关重要。尽管多模态大语言模型结合了视觉感知与自然语言,但现有医学定位流程仍依赖带空间提示的监督微调,难以应对临床实践中常见的隐含查询。本文提出统一医学推理定位(UMRG)任务,定义需结合临床推理与像素级定位的新范式。发布U-MRG-14K数据集,包含14,000个样本,涵盖10种模态、15个超类别和108个具体类别,提供像素级掩码、隐含临床问题及推理轨迹。提出MedReasoner框架,将推理与分割解耦:使用强化学习优化的MLLM推理器生成空间提示,冻结的分割专家将其转为掩码,通过格式与准确率奖励实现对齐。该方法在U-MRG-14K上达到最先进性能,并展现出对未见临床查询的强大泛化能力,证明强化学习在可解释医学定位中的巨大潜力。
原文摘要 · Abstract (English)
Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding pipelines still rely on supervised fine-tuning with explicit spatial hints, making them ill-equipped to handle the implicit queries common in clinical practice. This work makes three core contributions. We first define Unified Medical Reasoning Grounding (UMRG), a novel vision-language task that demands clinical reasoning and pixel-level grounding. Second, we release U-MRG-14K, a dataset of 14K samples featuring pixel-level masks alongside implicit clinical queries and reasoning traces, spanning 10 modalities, 15 super-categories, and 108 specific categories. Finally, we introduce MedReasoner, a modular framework that distinctly separates reasoning from segmentation: an MLLM reasoner is optimized with reinforcement learning, while a frozen segmentation expert converts spatial prompts into masks, with alignment achieved through format and accuracy rewards. MedReasoner achieves state-of-the-art performance on U-MRG-14K and demonstrates strong generalization to unseen clinical queries, underscoring the significant promise of reinforcement learning for interpretable medical grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。