提出首个统一框架评估CT报告生成指标的临床有效性。
CTest-Metric: A Unified Framework to Assess Clinical Validity of Metrics for CT Report Generation
- 设计三模块测试框架:风格泛化、合成错误注入、专家评分相关性。
- 发现GREEN Score最贴近专家判断,而CRG反而与专家意见负相关。
- 适合医学AI开发者与临床研究者用于指标筛选与验证。
在生成式AI时代,尽管关键医疗任务日益自动化,但放射科报告生成(RRG)仍依赖欠优的评估指标。由于缺乏统一、明确的框架来评估这些指标在临床场景中的鲁棒性与适用性,领域专用指标的研发面临挑战。为此,我们提出CTest-Metric——首个统一的指标评估框架,包含三个模块:(i) 基于LLM的重述测试写作风格泛化性(WSG);(ii) 分级严重度的合成错误注入(SEI);(iii) 使用175例“分歧”病例的临床医生评分进行指标-专家相关性分析(MvE)。我们在七种基于CT-CLIP编码器的LLM上评估了八种常用指标(BLEU、ROUGE、METEOR、BERTScore-F1、F1-RadGraph、RaTEScore、GREEN Score、CRG)。结果表明,词汇类生成指标对风格变化高度敏感;GREEN Score与专家判断相关性最强(斯皮尔曼系数~0.70),而CRG呈现负相关;BERTScore-F1对事实性错误注入最不敏感。我们将开源框架、代码及部分匿名化评估数据(重述/错误注入后的CT报告),以支持可复现的基准测试与未来指标开发。
原文摘要 · Abstract (English)
In the generative AI era, where even critical medical tasks are increasingly automated, radiology report generation (RRG) continues to rely on suboptimal metrics for quality assessment. Developing domain-specific metrics has therefore been an active area of research, yet it remains challenging due to the lack of a unified, well-defined framework to assess their robustness and applicability in clinical contexts. To address this, we present CTest-Metric, a first unified metric assessment framework with three modules determining the clinical feasibility of metrics for CT RRG. The modules test: (i) Writing Style Generalizability (WSG) via LLM-based rephrasing; (ii) Synthetic Error Injection (SEI) at graded severities; and (iii) Metrics-vs-Expert correlation (MvE) using clinician ratings on 175 "disagreement" cases. Eight widely used metrics (BLEU, ROUGE, METEOR, BERTScore-F1, F1-RadGraph, RaTEScore, GREEN Score, CRG) are studied across seven LLMs built on a CT-CLIP encoder. Using our novel framework, we found that lexical NLG metrics are highly sensitive to stylistic variations; GREEN Score aligns best with expert judgments (Spearman~0.70), while CRG shows negative correlation; and BERTScore-F1 is least sensitive to factual error injection. We will release the framework, code, and allowable portion of the anonymized evaluation data (rephrased/error-injected CT reports), to facilitate reproducible benchmarking and future metric development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。