提出新评估指标CRG Score,更准确衡量放射科报告生成的临床正确性。
CRG Score: A Distribution-Aware Clinical Metric for Radiology Report Generation
- 基于参考报告中的临床异常,仅评估明确描述的内容。
- 通过平衡标签分布惩罚,避免对简单预测的偏倚。
- 适合作为临床对齐的奖励函数,适用于各类大模型评估。
长文本放射科报告生成的评估极具挑战性。传统自然语言生成指标无法捕捉临床正确性,而基于大模型的指标又缺乏泛化能力。临床准确性指标虽更相关,但易受类别不平衡影响,常偏向于简单预测。本文提出CRG Score,一种考虑标签分布、可自适应调整的评估指标,仅评估参考报告中明确提及的临床相关异常。该指标支持二值和结构化标签(如类型、位置),可与任意大模型结合用于特征提取。通过根据标签分布平衡惩罚力度,实现更公平、更稳健的评估,可作为临床对齐的奖励函数使用。
原文摘要 · Abstract (English)
Evaluating long-context radiology report generation is challenging. NLG metrics fail to capture clinical correctness, while LLM-based metrics often lack generalizability. Clinical accuracy metrics are more relevant but are sensitive to class imbalance, frequently favoring trivial predictions. We propose the CRG Score, a distribution-aware and adaptable metric that evaluates only clinically relevant abnormalities explicitly described in reference reports. CRG supports both binary and structured labels (e.g., type, location) and can be paired with any LLM for feature extraction. By balancing penalties based on label distribution, it enables fairer, more robust evaluation and serves as a clinically aligned reward function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。