用医生指导的LLM提升胸片报告生成评估精度
Ran Score: a LLM-based Evaluation Score for Radiology Report Generation
- 结合医生经验与大模型,从报告中提取多标签病灶
- 在多个数据集上将评估准确率提升至0.956,超基准15.7个百分点
- 特别擅长评估罕见异常,适合临床报告生成模型评测
胸片报告生成与自动评估受限于对低频异常识别不足及对否定、模糊等临床语言处理不当。本文构建一种结合临床专家知识与大语言模型的多标签病灶提取框架,并据此提出Ran Score,一种基于病灶层面的报告评估指标。基于三个非重叠的MIMIC-CXR-EN队列及一个独立的ChestX-CN验证队列,优化提示词、建立放射科医生标注的参考标准,并评估报告生成模型。优化后的框架在MIMIC-CXR-EN开发集上的宏平均得分从0.753提升至0.956,在可比标签上超越CheXbert基准15.7个百分点,且在ChestX-CN验证集上表现出良好泛化能力。结果表明,医生引导的提示优化能显著提高与放射科医生参考标准的一致性,而Ran Score可实现对报告真实性的病灶级评估,尤其适用于低频异常。
原文摘要 · Abstract (English)
Chest X-ray report generation and automated evaluation are limited by poor recognition of low-prevalence abnormalities and inadequate handling of clinically important language, including negation and ambiguity. We develop a clinician-guided framework combining human expertise and large language models for multi-label finding extraction from free-text chest X-ray reports and use it to define Ran Score, a finding-level metric for report evaluation. Using three non-overlapping MIMIC-CXR-EN cohorts from a public chest X-ray dataset and an independent ChestX-CN validation cohort, we optimize prompts, establish radiologist-derived reference labels and evaluate report generation models. The optimized framework improves the macro-averaged score from 0.753 to 0.956 on the MIMIC-CXR-EN development cohort, exceeds the CheXbert benchmark by 15.7 percentage points on directly comparable labels, and shows robust generalization on the ChestX-CN validation cohort. Here we show that clinician-guided prompt optimization improves agreement with a radiologist-derived reference standard and that Ran Score enables finding-level evaluation of report fidelity, particularly for low-prevalence abnormalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。