统一评估放射科报告的开源框架,支持多种评分方式。
RadEval: A framework for radiology text evaluation
- 整合传统、语境和临床概念等多类评估指标
- 在多个数据集上实现强零样本检索性能
- 适合医疗AI研究者用于报告生成模型评测
我们提出RadEval,一个统一的开源框架,用于评估放射科文本。该框架整合了从经典n-gram重叠(BLEU、ROUGE)到上下文感知度量(BERTScore),再到基于临床概念的分数(F1CheXbert、F1RadGraph、RaTEScore、SRR-BERT、TemporalEntityF1)以及先进的大模型评估器(GREEN)。我们优化并标准化了实现,将GREEN扩展至支持多种影像模态,并采用更轻量级模型;同时预训练了一个领域专用的放射科编码器,在零样本检索任务中表现优异。我们还发布了一个丰富标注的专家数据集,包含超过450个临床显著错误标签,并展示了不同指标与放射科医生判断的相关性。最后,RadEval提供统计检验工具和跨多个公开数据集的基线模型评估,推动放射科报告生成领域的可复现性和稳健基准测试。
原文摘要 · Abstract (English)
We introduce RadEval, a unified, open-source framework for evaluating radiology texts. RadEval consolidates a diverse range of metrics, from classic n-gram overlap (BLEU, ROUGE) and contextual measures (BERTScore) to clinical concept-based scores (F1CheXbert, F1RadGraph, RaTEScore, SRR-BERT, TemporalEntityF1) and advanced LLM-based evaluators (GREEN). We refine and standardize implementations, extend GREEN to support multiple imaging modalities with a more lightweight model, and pretrain a domain-specific radiology encoder, demonstrating strong zero-shot retrieval performance. We also release a richly annotated expert dataset with over 450 clinically significant error labels and show how different metrics correlate with radiologist judgment. Finally, RadEval provides statistical testing tools and baseline model evaluations across multiple publicly available datasets, facilitating reproducibility and robust benchmarking in radiology report generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。