用奖励模型实现可解释的放射科报告生成评估
ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
- 基于奖励模型与自定义评分标准,支持细粒度评分
- 与人工判断相关性更高,优于传统指标
- 适合需要可解释评估的医疗AI研究者
自动化放射科报告生成(R2Gen)已取得显著进展,但其评估仍面临挑战。传统指标依赖机械匹配或仅关注病灶实体,难以与人工评价一致。为此,我们提出ReFINE,一种专为R2Gen设计的自动评估框架。该框架利用奖励模型,结合基于边距的奖励强化损失,并通过定制化训练数据设计,支持用户自定义评估标准。借助GPT-4,我们构建了高效的数据生成流程,基于两种不同评分体系生成大量报告样本及对应分数。通过配对规则将高质量与低质量报告配对,训练大模型输出细粒度奖励。奖励控制损失使模型同时输出多项独立评分,总和即为最终得分。实验表明,ReFINE与人类判断相关性显著提升,且在模型选择中表现更优。系统不仅提供总体分,还输出各项子分,增强可解释性,并支持跨多种评估体系灵活训练。
原文摘要 · Abstract (English)
Automated radiology report generation (R2Gen) has advanced significantly, introducing challenges in accurate evaluation due to its complexity. Traditional metrics often fall short by relying on rigid word-matching or focusing only on pathological entities, leading to inconsistencies with human assessments. To bridge this gap, we introduce ReFINE, an automatic evaluation metric designed specifically for R2Gen. Our metric utilizes a reward model, guided by our margin-based reward enforcement loss, along with a tailored training data design that enables customization of evaluation criteria to suit user-defined needs. It not only scores reports according to user-specified criteria but also provides detailed sub-scores, enhancing interpretability and allowing users to adjust the criteria between different aspects of reports. Leveraging GPT-4, we designed an easy-to-use data generation pipeline, enabling us to produce extensive training data based on two distinct scoring systems, each containing reports of varying quality along with corresponding scores. These GPT-generated reports are then paired as accepted and rejected samples through our pairing rule to train an LLM towards our fine-grained reward model, which assigns higher rewards to the report with high quality. Our reward-control loss enables this model to simultaneously output multiple individual rewards corresponding to the number of evaluation criteria, with their summation as our final ReFINE. Our experiments demonstrate ReFINE's heightened correlation with human judgments and superior performance in model selection compared to traditional metrics. Notably, our model provides both an overall score and individual scores for each evaluation item, enhancing interpretability. We also demonstrate its flexible training across various evaluation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。