CLEAR用临床属性评估放射科报告,更准更懂医生想法。
CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation
- 构建表格框架,按五类关键属性逐项比对报告内容
- 在100份胸片报告上验证,与医生判断高度一致
- 适合评估报告生成模型的临床质量,尤其关注细节描述
现有评价指标往往缺乏细粒度和可解释性,难以捕捉候选报告与标准报告之间的细微临床差异。我们提出一种基于临床场景的表格化评估框架CLEAR,结合专家标注标签与属性级对比,不仅判断报告是否准确识别疾病存在与否,还评估其对阳性病灶在五个关键属性上的描述精度:首次出现、变化、严重程度、解剖位置及处理建议。相比以往方法,CLEAR的多维度属性级输出能更全面、更贴近临床地评估报告质量。为验证其临床相关性,我们联合五位认证放射科医生,构建CLEAR-Bench数据集,包含100份来自MIMIC-CXR的胸片报告,涵盖6个预定义属性与13种CheXpert疾病类别。实验表明,CLEAR在提取临床属性方面表现优异,其自动化指标与临床判断高度一致。
原文摘要 · Abstract (English)
Existing metrics often lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports, resulting in suboptimal evaluation. We introduce a Clinically-grounded tabular framework with Expert-curated labels and Attribute-level comparison for Radiology report evaluation (CLEAR). CLEAR not only examines whether a report can accurately identify the presence or absence of medical conditions, but also assesses whether it can precisely describe each positively identified condition across five key attributes: first occurrence, change, severity, descriptive location, and recommendation. Compared to prior works, CLEAR's multi-dimensional, attribute-level outputs enable a more comprehensive and clinically interpretable evaluation of report quality. Additionally, to measure the clinical alignment of CLEAR, we collaborate with five board-certified radiologists to develop CLEAR-Bench, a dataset of 100 chest X-ray reports from MIMIC-CXR, annotated across 6 curated attributes and 13 CheXpert conditions. Our experiments show that CLEAR achieves high accuracy in extracting clinical attributes and provides automated metrics that are strongly aligned with clinical judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。