arXiv:2512.10691cs.AIcs.CV2025-12被引 2

用强化学习提升医学影像报告生成与视觉定位效果

Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning

  • 结合强化学习与临床反馈奖励,优化报告生成与视觉定位
  • 在多个数据集上实现最优性能,显著超越基础微调模型
  • 适合医疗AI研究者及希望提升模型临床实用性的开发者

近期视觉-语言模型(VLM)在胸部X光(CXR)解读方面取得进展,但多数依赖监督微调(SFT),仅优化下一个词预测,未评估答案质量。相比之下,强化学习(RL)可融入任务特定反馈,其与显式中间推理(“思考”)结合已在可验证数学和编程任务中表现优异。为探究RL与思考在CXR VLM中的作用,我们基于Qwen3-VL对大量CXR数据进行大规模SFT以构建更新版RadVLM,随后通过冷启动SFT阶段赋予模型基本推理能力。接着采用组相对策略优化(GRPO)并使用临床相关的任务特定奖励进行报告生成与视觉定位优化,并在领域特定与通用领域版本的Qwen3-VL上进行匹配的强化学习实验,含与不含思考两种设置。结果表明,尽管强SFT仍是高性能的基础,但强化学习在两项任务上均带来额外提升,而显式思考并未进一步改善结果。在统一评估流程下,经过强化学习优化的RadVLM模型优于基线,达到报告生成与视觉定位的最新水平,凸显临床对齐的强化学习是医疗VLM中监督微调的有力补充。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have improved Chest X-ray (CXR) interpretation in multiple aspects. However, many medical VLMs rely solely on supervised fine-tuning (SFT), which optimizes next-token prediction without evaluating answer quality. In contrast, reinforcement learning (RL) can incorporate task-specific feedback, and its combination with explicit intermediate reasoning ("thinking") has demonstrated substantial gains on verifiable math and coding tasks. To investigate the effects of RL and thinking in a CXR VLM, we perform large-scale SFT on CXR data to build an updated RadVLM based on Qwen3-VL, followed by a cold-start SFT stage that equips the model with basic thinking ability. We then apply Group Relative Policy Optimization (GRPO) with clinically grounded, task-specific rewards for report generation and visual grounding, and run matched RL experiments on both domain-specific and general-domain Qwen3-VL variants, with and without thinking. Across these settings, we find that while strong SFT remains crucial for high base performance, RL provides additional gains on both tasks, whereas explicit thinking does not appear to further improve results. Under a unified evaluation pipeline, the RL-optimized RadVLM models outperform their baseline counterparts and reach state-of-the-art performance on both report generation and grounding, highlighting clinically aligned RL as a powerful complement to SFT for medical VLMs.

医学影像强化学习报告生成视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。