用临床反馈强化学习,让胸片报告生成更准确
EditGRPO: Reinforcement Learning with Post-Rollout Edits for Clinically Accurate Chest X-Ray Report Generation
- 结合在线探索与离线指导,训练时注入逐句修正
- 在4个数据集上临床指标平均提升3.4%
- 对未见数据集泛化能力更强,平均增益5.9%
放射科报告生成需要先进的医学图像分析、有效的时间推理和精准的文本生成。尽管近期多模态大模型取得了进展,但其监督微调目标未明确对齐临床效果。本文提出EditGRPO,一种混合策略强化学习算法,专为通过临床激励优化生成而设计。该方法在训练采样过程中注入句子级详细修正,融合了在线探索与离线引导,缓解了强化学习中的探索困境与采样效率问题。应用于Qwen2.5-VL-3B模型,在四个主要数据集上平均临床指标提升3.4%,并在未见数据集上表现出更强泛化能力,平均性能提升5.9%。
原文摘要 · Abstract (English)
Radiology report generation requires advanced medical image analysis, effective temporal reasoning, and accurate text generation. Although recent innovations, particularly multimodal large language models, have shown improved performance, their supervised fine-tuning (SFT) objective is not explicitly aligned with clinical efficacy. In this work, we introduce EditGRPO, a mixed-policy reinforcement learning algorithm designed specifically to optimize the generation through clinically motivated rewards. EditGRPO integrates on-policy exploration with off-policy guidance by injecting sentence-level detailed corrections during training rollouts. This mixed-policy approach addresses the exploration dilemma and sampling efficiency issues typically encountered in RL. Applied to a Qwen2.5-VL-3B, EditGRPO outperforms both SFT and vanilla GRPO baselines, achieving an average improvement of 3.4\% in clinical metrics across four major datasets. Notably, EditGRPO also demonstrates superior out-of-domain generalization, with an average performance gain of 5.9\% on unseen datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。