揭示强化学习在医疗视觉语言模型中的真实作用,发现其主要提升已有能力的精准度。
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
- 通过控制实验分离视觉、监督微调与强化学习三者影响
- 强化学习在模型已有一定推理能力时最有效,提升准确率和采样效率
- 适合追求高效医疗多模态模型性能优化的研究者
强化学习(RL)被越来越多地用于后训练医疗视觉语言模型(VLMs),但尚不清楚其是否真正提升了医学视觉推理能力,还是仅强化了由监督微调(SFT)带来的行为。本研究通过受控实验,在三个维度上解耦这些影响:视觉、SFT 和 RL。利用 MedMNIST 作为多模态测试平台,我们通过对比视觉仅模型基准来评估视觉感知能力,通过 Accuracy@1 与 Pass@K 的对比量化推理支持与采样效率,并分析强化学习如何弥合支持差距及其跨模态迁移效果。结果表明,当模型已有非平凡推理支持(高 Pass@K)时,强化学习最为有效:它主要优化输出分布,提升 Acc@1 和采样效率;而 SFT 扩展了推理支持范围,使强化学习得以发挥作用。基于此,我们提出一种边界感知的训练方案,并在小规模、平衡的 PMC 多选题 VQA 子集上对 OctoMed 初始化模型进行强化学习后训练,实现了在六个医疗 VQA 基准上的优异平均表现。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is increasingly used to post-train medical Vision-Language Models (VLMs), yet it remains unclear whether RL improves medical visual reasoning or mainly sharpens behaviors already induced by supervised fine-tuning (SFT). We present a controlled study that disentangles these effects along three axes: vision, SFT, and RL. Using MedMNIST as a multi-modality testbed, we probe visual perception by benchmarking VLM vision towers against vision-only baselines, quantify reasoning support and sampling efficiency via Accuracy@1 versus Pass@K, and evaluate when RL closes the support gap and how gains transfer across modalities. We find that RL is most effective when the model already has non-trivial support (high Pass@K): it primarily sharpens the output distribution, improving Acc@1 and sampling efficiency, while SFT expands support and makes RL effective. Based on these findings, we propose a boundary-aware recipe and instantiate it by RL post-training an OctoMed-initialized model on a small, balanced subset of PMC multiple-choice VQA, achieving strong average performance across six medical VQA benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。