arXiv:2504.04336cs.CLcs.AI2025-04被引 30

用合成数据训练大模型,自动发现放射科报告中的错误。

Generative Large Language Models Trained for Detecting Errors in Radiology Reports

  • 用GPT-4生成带错/无错的胸片报告,构建双源数据集。
  • 微调后的Llama-3在四类错误上平均F1达0.78,最高达0.828。
  • 真实医生评估确认模型可有效识别临床实际错误。

本回顾性研究构建了包含两部分的数据集:第一部分为1,656份由GPT-4根据指定提示生成的合成胸片报告,其中828份无错误,828份含错误;第二部分包含614份报告:307份2011–2016年MIMIC-CXR数据库中无错误的真实报告,以及对应生成的307份含错合成报告。所有错误分为四类:否定、左右混淆、时间间隔变化及转录错误。采用零样本提示、少样本提示或微调策略,对Llama-3、GPT-4和BiomedBERT等模型进行优化。通过F1分数、95%置信区间及配对t检验评估性能,并由放射科医生进一步验证。零样本提示下,微调后的Llama-3-70B-Instruct表现最佳,各项错误检测的F1分别为:否定错误0.769,左右错误0.772,时间间隔错误0.750,转录错误0.828,整体0.780。真实评估中,两名放射科医生审查200份模型输出报告,99份被双方确认含错,163份至少一名医生确认含错。表明基于合成与真实数据微调的生成式大模型显著提升放射科报告错误检测能力。

原文摘要 · Abstract (English)

In this retrospective study, a dataset was constructed with two parts. The first part included 1,656 synthetic chest radiology reports generated by GPT-4 using specified prompts, with 828 being error-free synthetic reports and 828 containing errors. The second part included 614 reports: 307 error-free reports between 2011 and 2016 from the MIMIC-CXR database and 307 corresponding synthetic reports with errors generated by GPT-4 on the basis of these MIMIC-CXR reports and specified prompts. All errors were categorized into four types: negation, left/right, interval change, and transcription errors. Then, several models, including Llama-3, GPT-4, and BiomedBERT, were refined using zero-shot prompting, few-shot prompting, or fine-tuning strategies. Finally, the performance of these models was evaluated using the F1 score, 95\% confidence interval (CI) and paired-sample t-tests on our constructed dataset, with the prediction results further assessed by radiologists. Using zero-shot prompting, the fine-tuned Llama-3-70B-Instruct model achieved the best performance with the following F1 scores: 0.769 for negation errors, 0.772 for left/right errors, 0.750 for interval change errors, 0.828 for transcription errors, and 0.780 overall. In the real-world evaluation phase, two radiologists reviewed 200 randomly selected reports output by the model. Of these, 99 were confirmed to contain errors detected by the models by both radiologists, and 163 were confirmed to contain model-detected errors by at least one radiologist. Generative LLMs, fine-tuned on synthetic and MIMIC-CXR radiology reports, greatly enhanced error detection in radiology reports.

医学AI错误检测大模型放射科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。