arXiv:2603.06183cs.CLcs.AI2026-03被引 5

用临床真实场景评估肺部X光报告生成质量,更准更可信。

CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation

  • 基于临床上下文和指南构建评分框架,区分严重错误与无关细节。
  • 在多个专家标注数据集上与放射科医生判断高度一致(Kendall tau达0.71)。
  • 适合评估医学生成报告的可靠性,尤其关注诊断正确性和患者安全。

我们提出CRIMSON,一种基于临床背景的生成式胸部X光报告评估框架,从诊断准确性、上下文相关性和患者安全性三方面进行评价。不同于以往指标,CRIMSON整合了患者年龄、检查指征及指南决策规则等完整临床信息,避免正常或无意义发现对总分造成过度影响。该框架将错误分为全面的分类体系,涵盖假阳性、漏诊及八类属性级错误(如位置、严重程度、测量值、过度解读等)。每项发现依据与心血管胸科放射科医师合作制定的指南,赋予临床重要性等级(紧急、可行动非紧急、不可行动、预期/良性),实现基于严重性的加权,优先关注关键失误。通过六位认证放射科医生在ReXVal上的标注,验证了强一致性(肯德尔tau=0.61–0.71;皮尔逊r=0.71–0.84),并引入两个新基准:RadJudge(针对临床挑战性案例的通过/不通过测试)和RadPref(超过100对比较案例,含结构化错误分类、严重性建模及三位放射科医生1–5分质量评分)。在两者中,CRIMSON均与专家偏好最强匹配。项目代码、评估基准与微调后的MedGemma模型已开源,地址为https://github.com/rajpurkarlab/CRIMSON。

原文摘要 · Abstract (English)

We introduce CRIMSON, a clinically grounded evaluation framework for chest X-ray report generation that assesses reports based on diagnostic correctness, contextual relevance, and patient safety. Unlike prior metrics, CRIMSON incorporates full clinical context, including patient age, indication, and guideline-based decision rules, and prevents normal or clinically insignificant findings from exerting disproportionate influence on the overall score. The framework categorizes errors into a comprehensive taxonomy covering false findings, missing findings, and eight attribute-level errors (e.g., location, severity, measurement, and diagnostic overinterpretation). Each finding is assigned a clinical significance level (urgent, actionable non-urgent, non-actionable, or expected/benign), based on a guideline developed in collaboration with attending cardiothoracic radiologists, enabling severity-aware weighting that prioritizes clinically consequential mistakes over benign discrepancies. CRIMSON is validated through strong alignment with clinically significant error counts annotated by six board-certified radiologists in ReXVal (Kendalls tau = 0.61-0.71; Pearsons r = 0.71-0.84), and through two additional benchmarks that we introduce. In RadJudge, a targeted suite of clinically challenging pass-fail scenarios, CRIMSON shows consistent agreement with expert judgment. In RadPref, a larger radiologist preference benchmark of over 100 pairwise cases with structured error categorization, severity modeling, and 1-5 overall quality ratings from three cardiothoracic radiologists, CRIMSON achieves the strongest alignment with radiologist preferences. We release the metric, the evaluation benchmarks, RadJudge and RadPref, and a fine-tuned MedGemma model to enable reproducible evaluation of report generation, all available at https://github.com/rajpurkarlab/CRIMSON.

医学报告生成临床评估大模型评测放射科AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。