用轻量级专家指导强化学习,少数据生成更准确的放射科报告。
OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- 通过轻量级专家步骤将探索失败转化为偏好监督,实现单阶段强化学习训练
- 在仅需少量数据的条件下,于CheXpert Plus数据集上达到0.341的F1分数
- 适合资源有限但需高准确率报告生成的临床场景
放射科报告生成(RRG)旨在从胸部X光片自动生成临床可信的报告。现有方法通常依赖大规模配对语料和超大模型进行多阶段训练,导致数据与计算成本极高。本文提出Oracle-educated GRPO(OraPO),结合基于FactScore的奖励(FactS),在资源受限条件下解决RRG任务。OraPO通过轻量级专家步骤,将GRPO在罕见或困难病例上的探索失败转化为直接偏好监督,实现单阶段纯强化学习训练。FactS通过提取原子临床事实并验证其与真实标签的蕴含关系,提供密集且可解释的句子级奖励,使学习过程基于诊断证据。OraPO与FactS共同构建了一个紧凑高效的框架,在仅使用小规模基础视觉语言模型和普通硬件的情况下,以2-3个数量级更少的训练数据,实现了在CheXpert Plus数据集上的新SOTA性能(F1=0.341)。
原文摘要 · Abstract (English)
Radiology report generation (RRG) aims to automatically produce clinically faithful reports from chest X-ray images. Prevailing work typically follows a scale-driven paradigm, by multi-stage training over large paired corpora and oversized backbones, making pipelines highly data- and compute-intensive. In this paper, we propose Oracle-educated GRPO (OraPO) with a FactScore-based reward (FactS) to tackle the RRG task under constrained budgets. OraPO enables single-stage, RL-only training by converting failed GRPO explorations on rare or difficult studies into direct preference supervision via a lightweight oracle step. FactS grounds learning in diagnostic evidence by extracting atomic clinical facts and checking entailment against ground-truth labels, yielding dense, interpretable sentence-level rewards. Together, OraPO and FactS create a compact and powerful framework that significantly improves learning efficiency on clinically challenging cases, setting the new SOTA performance on the CheXpert Plus dataset (0.341 in F1) with 2--3 orders of magnitude less training data using a small base VLM on modest hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。