arXiv:2506.00711cs.LGcs.AI2025-06NeurIPS被引 42

首个综合医学影像、时序数据和文本的临床大模型,用新强化学习方法提升诊断准确率。

QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training

论文配图:QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
图 1 · 摘自论文原文
  • 采用领域感知的奖励优化策略,按疾病罕见度和模态难度动态调整训练奖励。
  • 在9个临床领域上平均诊断准确率提升43%,分割定位效果达OpenAI o4-mini水平。
  • 开源全部模型权重与训练过程,支持医疗AI研究复现与应用开发。

临床决策常需融合异构数据,但现有多模态语言模型仍以视觉为主,难以跨专科泛化。我们提出QoQ-Med-7B/32B,首个开放通用的临床基础模型,能联合推理医学图像、时序信号与文本报告。模型采用领域感知相对策略优化(DRPO)训练,通过层级化归一化奖励,根据领域稀有性和模态难度动态调整,缓解数据分布偏斜导致的性能失衡。在涵盖9个临床领域的261万条指令微调数据上训练,相比无评判器的GRPO等方法,整体视觉领域宏F1平均提升43%。此外,基于重症分割数据训练后,其可精准定位病灶区域,交并比(IoU)为开源模型的10倍,达到OpenAI o4-mini水平。为促进可复现性与下游研究,我们公开(i)完整模型权重,(ii)模块化训练流程,(iii)所有中间推理轨迹,详见https://github.com/DDVD233/QoQ_Med。

原文摘要 · Abstract (English)

Clinical decision-making routinely demands reasoning over heterogeneous data, yet existing multimodal language models (MLLMs) remain largely vision-centric and fail to generalize across clinical specialties. To bridge this gap, we introduce QoQ-Med-7B/32B, the first open generalist clinical foundation model that jointly reasons across medical images, time-series signals, and text reports. QoQ-Med is trained with Domain-aware Relative Policy Optimization (DRPO), a novel reinforcement-learning objective that hierarchically scales normalized rewards according to domain rarity and modality difficulty, mitigating performance imbalance caused by skewed clinical data distributions. Trained on 2.61 million instruction tuning pairs spanning 9 clinical domains, we show that DRPO training boosts diagnostic performance by 43% in macro-F1 on average across all visual domains as compared to other critic-free training methods like GRPO. Furthermore, with QoQ-Med trained on intensive segmentation data, it is able to highlight salient regions related to the diagnosis, with an IoU 10x higher than open models while reaching the performance of OpenAI o4-mini. To foster reproducibility and downstream research, we release (i) the full model weights, (ii) the modular training pipeline, and (iii) all intermediate reasoning traces at https://github.com/DDVD233/QoQ_Med.

临床大模型多模态强化学习医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。