三阶段AI框架大幅提升放射科报告纠错准确率,成本降四成。
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
- 分三步走:提取可疑项、检测错误、剔除假阳性。
- 纠错准确率提升至15.9%,每千份报告成本降至5.58美元。
- 适合医疗AI质检团队,尤其关注高精度低耗的场景。
背景:由于错误发生率低,基于大语言模型(LLM)的报告校对正向预测值(PPV)有限。目的:评估三阶段LLM框架是否相比基线方法提升PPV并降低运营成本。方法:对来自MIMIC-III数据库的1,000份连续放射科报告(每类250份:摄影、超声、CT、MRI)进行回顾性分析,另用CheXpert和Open-i两个外部数据集作为验证集。测试三种LLM框架:(1)单提示检测器;(2)提取器+检测器;(3)提取器+检测器+假阳性验证器。精确度以PPV和绝对真阳性率(aTPR)衡量。效率通过模型推理费用与人工审阅报酬计算。统计检验采用聚簇自助法、精确麦纳玛尔检验及霍尔姆-邦费罗尼校正。结果:框架的PPV从0.063(95% CI: 0.036–0.101,框架1)升至0.079(0.049–0.118,框架2),显著提高至0.159(0.090–0.252,框架3;与基线相比P<.001)。aTPR保持稳定(0.012–0.014;P≥.84)。每千份报告的运营成本从框架1的9.72美元降至框架3的5.58美元,降幅达42.6%;框架2为6.85美元,降幅18.5%。人工审阅量由192份降至88份。外部验证支持框架3在CheXpert上达到0.133的PPV、Open-i上0.105,aTPR稳定在0.007。结论:三阶段框架显著提升纠错准确率并降低运营成本,同时维持检测性能,为人工智能辅助放射科报告质量控制提供有效方案。
原文摘要 · Abstract (English)
Background: The positive predictive value (PPV) of large language model (LLM)-based proofreading for radiology reports is limited due to the low error prevalence. Purpose: To assess whether a three-pass LLM framework enhances PPV and reduces operational costs compared with baseline approaches. Materials and Methods: A retrospective analysis was performed on 1,000 consecutive radiology reports (250 each: radiography, ultrasonography, CT, MRI) from the MIMIC-III database. Two external datasets (CheXpert and Open-i) were validation sets. Three LLM frameworks were tested: (1) single-prompt detector; (2) extractor plus detector; and (3) extractor, detector, and false-positive verifier. Precision was measured by PPV and absolute true positive rate (aTPR). Efficiency was calculated from model inference charges and reviewer remuneration. Statistical significance was tested using cluster bootstrap, exact McNemar tests, and Holm-Bonferroni correction. Results: Framework PPV increased from 0.063 (95% CI, 0.036-0.101, Framework 1) to 0.079 (0.049-0.118, Framework 2), and significantly to 0.159 (0.090-0.252, Framework 3; P<.001 vs. baselines). aTPR remained stable (0.012-0.014; P>=.84). Operational costs per 1,000 reports dropped to USD 5.58 (Framework 3) from USD 9.72 (Framework 1) and USD 6.85 (Framework 2), reflecting reductions of 42.6% and 18.5%, respectively. Human-reviewed reports decreased from 192 to 88. External validation supported Framework 3's superior PPV (CheXpert 0.133, Open-i 0.105) and stable aTPR (0.007). Conclusion: A three-pass LLM framework significantly enhanced PPV and reduced operational costs, maintaining detection performance, providing an effective strategy for AI-assisted radiology report quality assurance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。