arXiv:2606.18797cs.CL2026-06

用大模型提升放射科报告临床评估精度,区分关键错误与无害表述。

Beyond Scalar Scores: Exploring LLM-based Metrics for Clinical Significance Evaluation in Radiology Reports

论文配图:Beyond Scalar Scores: Exploring LLM-based Metrics for Clinical Significance Evaluation in Radiology Reports
图 1 · 摘自论文原文
  • 基于大模型构建可解释的评估指标,精准识别临床错误。
  • 训练后指标超越32B级医学大模型,接近商用水平。
  • 单次推理更实用,双次推理未必更好,适合成本敏感场景。

生成放射科报告的可靠评估需确保临床准确性,遗漏关键发现或误判影像表现可能直接影响患者治疗。现有指标将报告质量简化为无医学依据的单一数值,掩盖了真实临床需求。尽管大语言模型(LLMs)具备丰富的医学知识,但仍难以区分临床显著错误与无害表述差异。本文以ReEvalMed基准测试为实验平台,从检测真实临床错误(“区分度”)和容忍无意义变化(“鲁棒性”)两个维度评估8个LLM评估器在单次与双次推理设置下的表现。结果发现普遍存在的区分偏差:模型能有效检测错误,却过度惩罚无害改写。为此,我们合成4000对报告,基于Qwen3-8B与MedGemma-4B训练轻量级可解释指标。该指标显著优化临床意义边界,超越32B规模医学大模型,并保持与专有模型相当的竞争力。值得注意的是,更昂贵的双次推理未稳定提升整体性能,主要在区分度与鲁棒性间做权衡。研究建议在成本敏感场景采用单次训练指标,双次推理仅用于对区分-鲁棒平衡要求极高的场合。数据集与指标将开源。

原文摘要 · Abstract (English)

Reliable evaluation of generated radiology reports requires strict clinical accuracy, as omitted critical findings or mischaracterized radiographic observations can directly affect patient care. Existing metrics obscure this requirement by reducing report quality to a medically ungrounded scalar. Although Large Language Models (LLMs) possess rich medical knowledge, they likewise struggle to draw a reliable boundary between clinically significant errors and harmless variation. We study this boundary using ReEvalMed benchmark as testbed and evaluate metric-level clinical significance from detecting true clinical errors ("Discrimination") and tolerating insignificant variations ("Robustness"). Across 8 LLM evaluators under one-pass and two-pass settings, we identify a widespread discrimination bias: models effectively detect errors but also over-penalize harmless rephrasings. To mitigate this, we synthesize 4k report pairs and train lightweight interpretable metrics on Qwen3-8B and MedGemma-4B. Our trained metric sharpens the clinical significance boundary, surpassing 32B-scale medical LLMs and remaining competitive with proprietary models. Crucially, the more costly two-pass setting fails to consistently improve overall performance and mainly trades discrimination for robustness. These findings suggest one-pass trained metrics as the practical choice for cost-sensitive deployment, with two-pass inference reserved for settings where D-R balance is critical. We will release the dataset and metric.

医学AI大模型评估放射科报告可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。