arXiv:2503.21004cs.CL2025-03被引 1

大模型可精准自动提取肺栓塞报告关键信息,效果媲美人工。

Evaluating Large Language Models for Automated Clinical Abstraction in Pulmonary Embolism Registries: Performance Across Model Sizes, Versions, and Parameters

  • 用不同规模的大模型自动抽取肺栓塞影像报告中的医学概念。
  • 70亿参数以上模型准确率超96%,双模型校验后准确率仍超95%。
  • 适合医疗数据自动化处理,减少人工标注负担。

肺栓塞(PE)登记系统推动临床研究进展,但依赖耗时的人工提取放射科报告。我们评估了公开可用的大语言模型(LLM)能否在不降低数据质量的前提下,自动从CT肺栓塞(CTPE)报告中提取关键概念。测试了四种Llama-3变体(3.0 8B、3.1 8B、3.1 70B、3.3 70B)及Phi-4(P4)14B、Gemma-3 27B(G3)共六种模型,分别在MIMIC-IV和杜克大学的各250份双标注CTPE报告上评估。结果以准确率、阳性预测值(PPV)、阴性预测值(NPV)与人工金标准对比,涵盖模型规模、温度设置和样本数量。所有概念平均准确率随模型规模提升:L3-0 8B为0.83,L3-1 8B为0.91,两个70B版本均为0.96;P4 14B达0.98,G3表现相当。跨数据集差异小于0.03,体现良好外部鲁棒性。双模型一致性分析(L3 70B + P4 14B)显示:肺栓塞存在性PPV ≥ 0.95,NPV ≥ 0.98;位置、血栓负荷、右心室受累、图像伪影等均保持PPV ≥ 0.90,NPV ≥ 0.95。单个概念标注分歧少于4%,超过75%报告实现完全一致。G3表现相当。因此,大模型可提供可扩展、高精度的肺栓塞登记自动化方案,双模型审查流程可极小化人工干预保障数据质量。

原文摘要 · Abstract (English)

Pulmonary embolism (PE) registries accelerate practice-improving research but depend on resource-intensive manual abstraction of radiology reports. We evaluated whether openly available large-language models (LLMs) can automate concept extraction from computed-tomography PE (CTPE) reports without sacrificing data quality. Four Llama-3 (L3) variants (3.0 8 B, 3.1 8 B, 3.1 70 B, 3.3 70 B) and two reviewer models Phi-4 (P4) 14 B and Gemma-3 27 B (G3) were tested on 250 dual-annotated CTPE reports each from MIMIC-IV and Duke University. Outcomes were accuracy, positive predictive value (PPV), and negative predictive value (NPV) versus a human gold standard across model sizes, temperature settings, and shot counts. Mean accuracy across all concepts increased with scale: 0.83 (L3-0 8 B), 0.91 (L3-1 8 B), and 0.96 for both 70 B variants; P4 14 B achieved 0.98; G3 matched. Accuracy differed by < 0.03 between datasets, underscoring external robustness. In dual-model concordance analysis (L3 70 B + P4 14 B), PE-presence PPV was >= 0.95 and NPV >= 0.98, while location, thrombus burden, right-heart strain, and image-quality artifacts each maintained PPV >= 0.90 and NPV >= 0.95. Fewer than 4% of individual concept annotations were discordant, and complete agreement was observed in more than 75% of reports. G3 performed comparably. LLMs therefore offer a scalable, accurate solution for PE registry abstraction, and a dual-model review workflow can further safeguard data quality with minimal human oversight.

医学信息抽取大模型应用自动化标注肺栓塞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。