用自生成数据解决医学AI训练缺高质量标注的难题
MedGR$^2$: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning
- 构建生成器与奖励模型协同进化,自动产生物理-影像-文本多模态医学数据
- 用自生成数据微调后模型性能超越人工标注数据训练的基线
- 小模型经自生成数据训练后媲美十倍参数的大模型,适合资源有限场景
医学视觉语言模型应用受制于高质量专家标注数据稀缺。现有数据集上监督微调难以泛化到新模态和任务,而强化学习因缺乏可靠奖励信号受阻。为此,我们提出医学推理生成奖励学习框架MedGR²,通过协同开发数据生成器与奖励模型,实现高质、多模态医学数据的自动化持续生成,既可用于监督微调,也可用于强化学习。实验表明,使用MedGR²生成数据进行微调的模型已超越基于大规模人工标注数据训练的基线;进一步结合组相对策略优化(GRPO)进行强化学习,模型在跨模态与跨任务泛化上达到当前最优,显著优于专用强化学习方法。此外,该框架赋能的小模型性能可媲美参数量超10倍的基座模型。MedGR²为高风险领域提供了数据高效学习新范式,将数据短缺问题转化为数据生成能力,充分释放强化学习在构建通用医疗AI中的潜力。
原文摘要 · Abstract (English)
The application of Vision-Language Models (VLMs) in medicine is critically hampered by the scarcity of high-quality, expert-annotated data. Supervised Fine-Tuning (SFT) on existing datasets often leads to poor generalization on unseen modalities and tasks, while Reinforcement Learning (RL), a promising alternative, is stymied by the lack of reliable reward signals in this data-scarce domain. To break this impasse, we introduce Generative Reward Learning for Medical Reasoning (MedGR$^2$), a novel framework that creates a self-improving virtuous cycle. MedGR$^2$ co-develops a data generator and a reward model, enabling the automated, continuous creation of high-quality, multi-modal medical data that serves as both a superior training source for SFT and RL. Our experiments demonstrate that SFT with MedGR$^2$-produced data already surpasses baselines trained on large-scale, human-curated datasets. Crucially, when leveraging this data for RL via Group Relative Policy Optimization (GRPO), our model achieves state-of-the-art cross-modality and cross-task generalization, significantly outperforming specialized RL-based methods. Furthermore, our compact model, empowered by MedGR$^2$, achieves performance competitive with foundation models possessing over 10 times more parameters. MedGR$^2$ presents a new paradigm for data-efficient learning in high-stakes domains, transforming the problem from data scarcity to data generation and unlocking the full potential of RL for building truly generalizable medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。