arXiv:2510.00280cs.CL2025-10EMNLP被引 7

现有医学报告评估指标与临床实际判断脱节,新框架提升评估可信度。

ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment

  • 基于临床真实判断设计元评估框架,聚焦临床一致性与指标能力。
  • 发现现有指标无法区分重要错误、过度惩罚无害变化,且对严重程度不敏感。
  • 适用于医疗AI评估改进,尤其适合关注临床实用性的研究者。

自动生成的放射科报告虽在现有评估指标上得分高,却难以获得临床医生信任。这暴露出当前评估方法在理解临床语义上的根本缺陷。本文重新思考评估指标的设计,提出一种基于临床实践的元评估框架,定义了涵盖临床一致性及判别力、鲁棒性、单调性等关键能力的评价标准。利用一个细粒度标注数据集(包含真实报告与重写版本对),标注错误类型、临床重要性标签及解释,系统评估现有指标,揭示其在识别临床显著错误、处理无害变异、保持不同严重程度错误间一致性方面的不足。该框架为构建更符合临床需求的评估方法提供指导。

原文摘要 · Abstract (English)

Automatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians' trust. This gap reveals fundamental flaws in how current metrics assess the quality of generated reports. We rethink the design and evaluation of these metrics and propose a clinically grounded Meta-Evaluation framework. We define clinically grounded criteria spanning clinical alignment and key metric capabilities, including discrimination, robustness, and monotonicity. Using a fine-grained dataset of ground truth and rewritten report pairs annotated with error types, clinical significance labels, and explanations, we systematically evaluate existing metrics and reveal their limitations in interpreting clinical semantics, such as failing to distinguish clinically significant errors, over-penalizing harmless variations, and lacking consistency across error severity levels. Our framework offers guidance for building more clinically reliable evaluation methods.

医学报告评估指标临床对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。