arXiv:2504.09258cs.CVcs.MM2025-04被引 8

用强化学习提升病理图像理解,让AI诊断更可靠且可解释。

PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks

  • 基于强化学习设计推理流程,约束逻辑正确性与结果准确性
  • 在病理问答任务中准确率比基线高14%,参数量更小却性能更强
  • 跨模态迁移能力突出,多类医学影像泛化性能平均提升17.3%

病理图像诊断常受限于专家资源不足和区域差异,亟需借助视觉语言模型(VLMs)实现自动化。传统多模态模型重结果轻推理过程,影响临床决策可靠性。为此,我们提出专用于病理图像的PathVLM-R1模型,基于Qwen2.5-VL-7B-Instruct,并通过精细化后训练策略提升性能。首先利用病理数据进行监督微调,构建基础病理模型;随后引入分组相对策略优化(GRPO),设计双奖励驱动的强化学习优化机制,通过跨模态过程奖励与结果准确率奖励,严格约束推理逻辑与输出精度。在病理图像问答任务中,PathVLM-R1准确率较基线提升14%,优于参数量更大的Qwen2.5-VL-32B。在涵盖计算机断层扫描(CT)、皮肤镜、眼底摄影和光学相干断层扫描(OCT)四类医学影像的跨域评估中,其迁移性能相较传统SFT方法平均提升17.3%。结果表明,PathVLM-R1不仅提升诊断准确率,还具备强泛化能力与扩展潜力。

原文摘要 · Abstract (English)

The diagnosis of pathological images is often limited by expert availability and regional disparities, highlighting the importance of automated diagnosis using Vision-Language Models (VLMs). Traditional multimodal models typically emphasize outcomes over the reasoning process, compromising the reliability of clinical decisions. To address the weak reasoning abilities and lack of supervised processes in pathological VLMs, we have innovatively proposed PathVLM-R1, a visual language model designed specifically for pathological images. We have based our model on Qwen2.5-VL-7B-Instruct and enhanced its performance for pathological tasks through meticulously designed post-training strategies. Firstly, we conduct supervised fine-tuning guided by pathological data to imbue the model with foundational pathological knowledge, forming a new pathological base model. Subsequently, we introduce Group Relative Policy Optimization (GRPO) and propose a dual reward-driven reinforcement learning optimization, ensuring strict constraint on logical supervision of the reasoning process and accuracy of results via cross-modal process reward and outcome accuracy reward. In the pathological image question-answering tasks, the testing results of PathVLM-R1 demonstrate a 14% improvement in accuracy compared to baseline methods, and it demonstrated superior performance compared to the Qwen2.5-VL-32B version despite having a significantly smaller parameter size. Furthermore, in out-domain data evaluation involving four medical imaging modalities: Computed Tomography (CT), dermoscopy, fundus photography, and Optical Coherence Tomography (OCT) images: PathVLM-R1's transfer performance improved by an average of 17.3% compared to traditional SFT methods. These results clearly indicate that PathVLM-R1 not only enhances accuracy but also possesses broad applicability and expansion potential.

病理诊断视觉语言模型强化学习医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。