通过分层推理优化,让医学报告生成更准确且有依据。
HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize
- 分三层次优化推理、诊断与证据对齐
- 在两个数据集上显著降低临床幻觉
- 适合需要高可信度报告的医疗AI研究者
多模态大模型显著推进了放射科报告生成(RRG),但通过强化学习对齐时面临医学标注异质性挑战。传统组相对策略优化(GRPO)对整个生成过程赋予相同信用,导致片段干扰、标记稀释及证据与诊断脱节,加剧临床幻觉。我们提出HERO(分层证据推理优化),一种因子化策略优化框架,通过三个不同粒度对齐异质监督:分别在段落、标记和完成层级进行优化,并采用涵盖诊断准确性、推理质量与思考-答案一致性的异质奖励函数。在MIMIC-CXR和IU-Xray数据集上的实验表明,HERO优于强监督和强化学习基线,达到当前最优临床效能,生成更具证据支撑且思考-答案一致的报告,显著减少临床幻觉。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision. Vanilla Group Relative Policy Optimization (GRPO) assigns uniform credit across the entire generation, leading to segment interference, token dilution, and evidence--diagnosis decoupling, which exacerbates clinical hallucinations. We propose HERO (Hierarchical Evidential Reasoning Optimization), a factorized policy optimization framework that aligns heterogeneous supervision with three optimization granularities. HERO separately optimizes reasoning, diagnosis, and evidence grounding through complementary segment-, token-, and completion-level optimization with a heterogeneous reward formulation covering diagnostic accuracy, reasoning quality, and think--answer consistency. Experiments on MIMIC-CXR and IU-Xray show that HERO outperforms strong supervised and reinforcement learning baselines, achieving state-of-the-art clinical efficacy while producing more evidence-grounded and think--answer-consistent reports, thereby substantially mitigating clinical hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。