arXiv:2411.15122cs.CVcs.AI2024-11AAAI被引 30

首个公开的医学影像报告生成评估平台,统一评测标准。

ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation

  • 构建包含1万份影像的测试集,整合三大公开数据集
  • 采用8项指标评估模型表现,区分仅生成发现与完整报告能力
  • 推动医疗影像报告生成模型的公平对比,适合研究者与临床开发者

AI驱动的模型在胸部X光片报告自动生成方面展现出巨大潜力,但缺乏标准化评估基准。为此,我们推出ReXrank(https://rexrank.ai),一个面向AI辅助放射科报告生成的公开排行榜与挑战赛。该框架包含ReXGradient——目前最大的测试数据集,含10,000个影像研究案例,并整合了MIMIC-CXR、IU-Xray和CheXpert Plus三个公开数据集用于报告生成评估。ReXrank采用8项评价指标,分别评估仅生成发现部分与同时生成发现及结论部分的模型能力。通过提供标准化评估体系,ReXrank实现了模型性能的客观比较,并揭示其在多样化临床场景下的鲁棒性。未来可扩展至全谱医学影像自动报告评估。

原文摘要 · Abstract (English)

AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no standardized benchmark for objectively evaluating their performance. To address this, we present ReXrank, https://rexrank.ai, a public leaderboard and challenge for assessing AI-powered radiology report generation. Our framework incorporates ReXGradient, the largest test dataset consisting of 10,000 studies, and three public datasets (MIMIC-CXR, IU-Xray, CheXpert Plus) for report generation assessment. ReXrank employs 8 evaluation metrics and separately assesses models capable of generating only findings sections and those providing both findings and impressions sections. By providing this standardized evaluation framework, ReXrank enables meaningful comparisons of model performance and offers crucial insights into their robustness across diverse clinical settings. Beyond its current focus on chest X-rays, ReXrank's framework sets the stage for comprehensive evaluation of automated reporting across the full spectrum of medical imaging.

医学影像报告生成评估基准公开数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。