arXiv:2508.03733cs.LGcs.AI2025-08被引 3

首个实现胸部X光片分步推理的医学大模型,显著减少错误和冗长分析。

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning

  • 采用分阶段强化学习,让模型像医生一样边思考边回答。
  • 在14种疾病上召回率超现有模型,提升25.1%性能。
  • 适合临床辅助诊断、医学研究者及医疗AI开发者使用。

胸部X光(CXR)是临床最广泛应用的影像诊断手段之一,涵盖多种诊断任务。近年来,基于推理的多模态大语言模型(MLLM)被广泛用于提升诊断效率与可解释性。然而,现有模型多依赖一次性诊断方式,缺乏对推理过程的可验证监督,导致多任务诊断中存在推理过长、奖励稀疏和幻觉频发等问题。为此,我们提出CX-Mind,首个实现胸部X光任务中“思考-回答”交替推理的生成式模型,通过课程引导强化学习与可验证过程奖励(CuRL-VPR)驱动。构建了包含708,473张图像和2,619,148个样本的指令微调数据集CX-Set,生成42,828条由临床报告监督的高质量分步推理数据。在群组相对策略优化框架下分两阶段训练:先用封闭域任务稳定基础推理,再迁移至开放域诊断,并引入基于规则的条件过程奖励,避免依赖预训练奖励模型。大量实验表明,CX-Mind在视觉理解、文本生成和时空对齐方面显著优于现有医学与通用领域MLLM,相较同类专用于CXR的模型平均性能提升25.1%。在真实世界临床数据集Rui-CXR上,其14种疾病的平均召回率@1大幅超越第二名,多中心专家评估进一步证实其临床实用性。

原文摘要 · Abstract (English)

Chest X-ray (CXR) imaging is one of the most widely used diagnostic modalities in clinical practice, encompassing a broad spectrum of diagnostic tasks. Recent advancements have seen the extensive application of reasoning-based multimodal large language models (MLLMs) in medical imaging to enhance diagnostic efficiency and interpretability. However, existing multimodal models predominantly rely on "one-time" diagnostic approaches, lacking verifiable supervision of the reasoning process. This leads to challenges in multi-task CXR diagnosis, including lengthy reasoning, sparse rewards, and frequent hallucinations. To address these issues, we propose CX-Mind, the first generative model to achieve interleaved "think-answer" reasoning for CXR tasks, driven by curriculum-based reinforcement learning and verifiable process rewards (CuRL-VPR). Specifically, we constructed an instruction-tuning dataset, CX-Set, comprising 708,473 images and 2,619,148 samples, and generated 42,828 high-quality interleaved reasoning data points supervised by clinical reports. Optimization was conducted in two stages under the Group Relative Policy Optimization framework: initially stabilizing basic reasoning with closed-domain tasks, followed by transfer to open-domain diagnostics, incorporating rule-based conditional process rewards to bypass the need for pretrained reward models. Extensive experimental results demonstrate that CX-Mind significantly outperforms existing medical and general-domain MLLMs in visual understanding, text generation, and spatiotemporal alignment, achieving an average performance improvement of 25.1% over comparable CXR-specific models. On real-world clinical dataset (Rui-CXR), CX-Mind achieves a mean recall@1 across 14 diseases that substantially surpasses the second-best results, with multi-center expert evaluations further confirming its clinical utility across multiple dimensions.

医学AI多模态推理模型胸部X光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。