提出可解释的放射科报告评估框架,自动打细粒度分数并给出理由。
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
- 基于群体相对策略优化,按临床重要性动态调整错误类型权重。
- 在ReXVal上优于所有已有离线评估指标,接近GPT-4表现。
- 输出可读理由与子分,适合临床部署和医生审核。
自动评估放射科报告仍面临根本性挑战,源于缺乏临床可解释、细粒度的评估指标。现有方法或仅提供粗略总分,或依赖难以理解的黑箱模型,限制其在真实临床流程中的应用。我们提出RadReason,一种新型放射科报告评估框架,不仅能输出六类临床定义错误类型的细粒度子分,还能生成人类可读的解释,说明每项评分的依据。该方法基于组相对策略优化(Group Relative Policy Optimization),引入两项关键创新:(1) 子分动态加权,根据实时F1统计自适应优先处理临床挑战性错误类型;(2) 多数引导优势缩放,基于子分一致性推导提示难度以调整策略梯度更新。两者协同实现更稳定的优化,并更好对齐专家临床判断。在ReXVal基准上的实验表明,RadReason超越所有先前离线评估指标,达到与GPT-4评估相当的性能,同时保持可解释性、低成本,适合临床部署。代码将在发表后公开。
原文摘要 · Abstract (English)
Evaluating automatically generated radiology reports remains a fundamental challenge due to the lack of clinically grounded, interpretable, and fine-grained metrics. Existing methods either produce coarse overall scores or rely on opaque black-box models, limiting their usefulness in real-world clinical workflows. We introduce RadReason, a novel evaluation framework for radiology reports that not only outputs fine-grained sub-scores across six clinically defined error types, but also produces human-readable justifications that explain the rationale behind each score. Our method builds on Group Relative Policy Optimization and incorporates two key innovations: (1) Sub-score Dynamic Weighting, which adaptively prioritizes clinically challenging error types based on live F1 statistics; and (2) Majority-Guided Advantage Scaling, which adjusts policy gradient updates based on prompt difficulty derived from sub-score agreement. Together, these components enable more stable optimization and better alignment with expert clinical judgment. Experiments on the ReXVal benchmark show that RadReason surpasses all prior offline metrics and achieves parity with GPT-4-based evaluations, while remaining explainable, cost-efficient, and suitable for clinical deployment. Code will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。