arXiv:2509.15154cs.CV2025-09被引 1

通过伪标签与强化学习提升医疗多模态模型的事实准确性。

MedFact-R1: Towards Factual Medical Reasoning via Pseudo-Label Augmentation

  • 用伪标签监督微调引入外部医学知识
  • 强化学习优化后,事实准确率提升22.5%
  • 适合医疗AI可信推理研究者参考

确保事实一致性和可靠推理仍是医疗视觉-语言模型的核心挑战。我们提出MEDFACT-R1,一种两阶段框架,通过融合外部知识定位与强化学习来提升医疗事实推理能力。第一阶段采用伪标签监督微调(SFT)引入外部事实知识;第二阶段使用四类定制化事实奖励信号的组相对策略优化(GRPO),促进自我一致性推理。在三个公开医疗问答基准上,MEDFACT-R1相比先前最优方法,事实准确率最高提升22.5%。消融实验表明伪标签SFT冷启动的必要性,并验证了每项GRPO奖励的贡献,凸显知识定位与强化学习推理之间的协同作用。代码已开源:https://github.com/Garfieldgengliang/MEDFACT-R1。

原文摘要 · Abstract (English)

Ensuring factual consistency and reliable reasoning remains a critical challenge for medical vision-language models. We introduce MEDFACT-R1, a two-stage framework that integrates external knowledge grounding with reinforcement learning to improve the factual medical reasoning. The first stage uses pseudo-label supervised fine-tuning (SFT) to incorporate external factual expertise; while the second stage applies Group Relative Policy Optimization (GRPO) with four tailored factual reward signals to encourage self-consistent reasoning. Across three public medical QA benchmarks, MEDFACT-R1 delivers up to 22.5% absolute improvement in factual accuracy over previous state-of-the-art methods. Ablation studies highlight the necessity of pseudo-label SFT cold start and validate the contribution of each GRPO reward, underscoring the synergy between knowledge grounding and RL-driven reasoning for trustworthy medical AI. Codes are released at https://github.com/Garfieldgengliang/MEDFACT-R1.

医疗AI事实推理强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。