arXiv:2604.13598cs.LGstat.ME2026-04ACL被引 3

用证据感知奖励和自修正机制提升放射科报告生成的临床准确性

Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning

论文配图:Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning
图 1 · 摘自论文原文
  • 引入分组证据对齐奖励,强化真实阳性、补全漏诊、抑制误报
  • 通过LLM自动构建疾病感知偏好数据集,实现无监督自修正
  • 在两个胸片数据集上达到领先性能,适合医疗AI研发者参考

近期强化学习方法推动了放射科报告生成(RRG)的发展,但仍存在两大核心问题:(1)报告级奖励难以提供基于证据的临床指导;(2)现有方法缺乏显式的自我优化机制以对齐临床偏好。本文提出临床对齐的证据感知自修正强化学习(ESC-RL),包含两个关键组件。首先,分组证据对齐奖励(GEAR)提供分组级、证据感知反馈,强化真实阳性的一致性,恢复假阴性的漏诊,抑制假阳性的无依据内容。其次,自修正偏好学习(SPL)策略从多个噪声观测中自动构建可靠、疾病感知的偏好数据集,并利用大语言模型合成优化报告,无需人工标注。ESC-RL促进临床可信、疾病对齐的奖励信号,并支持训练过程中的持续自我改进。在两个公开胸片数据集上的大量实验表明,该方法持续提升性能并达到当前最优水平。

原文摘要 · Abstract (English)

Recent reinforcement learning (RL) approaches have advanced radiology report generation (RRG), yet two core limitations persist: (1) report-level rewards offer limited evidence-grounded guidance for clinical faithfulness; and (2) current methods lack an explicit self-improving mechanism to align with clinical preference. We introduce clinically aligned Evidence-aware Self-Correcting Reinforcement Learning (ESC-RL), comprising two key components. First, a Group-wise Evidence-aware Alignment Reward (GEAR) delivers group-wise, evidence-aware feedback. GEAR reinforces consistent grounding for true positives, recovers missed findings for false negatives, and suppresses unsupported content for false positives. Second, a Self-correcting Preference Learning (SPL) strategy automatically constructs a reliable, disease-aware preference dataset from multiple noisy observations and leverages an LLM to synthesize refined reports without human supervision. ESC-RL promotes clinically faithful, disease-aligned reward and supports continual self-improvement during training. Extensive experiments on two public chest X-ray datasets demonstrate consistent gains and state-of-the-art performance.

放射科报告强化学习自修正临床对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。