arXiv:2411.00044cs.CLcs.LG2024-11

用大模型自动标注肺栓塞,提升医疗数据研究效率

MIMIC-IV-Ext-PE: Using a large language model to predict pulmonary embolism phenotype in the MIMIC-IV dataset

  • 用微调的Bio_ClinicalBERT模型从影像报告中自动提取肺栓塞标签
  • 模型在19,942例患者中敏感性达92.4%,阳性预测值87.8%
  • 适用于需要大规模肺栓塞数据的临床研究与算法开发

肺栓塞(PE)是院内可预防死亡的主要原因。现有公开数据集缺乏足够的PE标注。本研究基于MIMIC-IV数据库,提取所有急诊和住院患者的CTPA影像报告,并由两名医生手动标注为急性肺栓塞阳性或阴性。随后,使用已微调的生物医学语言模型VTE-BERT进行自动标注。通过与人工判读对比验证了VTE-BERT的可靠性。结果表明,在全部19,942例急诊/入院患者中,该模型敏感性为92.4%,阳性预测值(PPV)为87.8%;而诊断代码在11,990名有出院诊断记录的住院患者中,敏感性为95.4%,PPV为83.8%。研究成功为近2万例CTPA添加标注,证实了半监督语言模型在血液病研究中的外部有效性。

原文摘要 · Abstract (English)

Pulmonary embolism (PE) is a leading cause of preventable in-hospital mortality. Advances in diagnosis, risk stratification, and prevention can improve outcomes. There are few large publicly available datasets that contain PE labels for research. Using the MIMIC-IV database, we extracted all available radiology reports of computed tomography pulmonary angiography (CTPA) scans and two physicians manually labeled the results as PE positive (acute PE) or PE negative. We then applied a previously finetuned Bio_ClinicalBERT transformer language model, VTE-BERT, to extract labels automatically. We verified VTE-BERT's reliability by measuring its performance against manual adjudication. We also compared the performance of VTE-BERT to diagnosis codes. We found that VTE-BERT has a sensitivity of 92.4% and positive predictive value (PPV) of 87.8% on all 19,942 patients with CTPA radiology reports from the emergency room and/or hospital admission. In contrast, diagnosis codes have a sensitivity of 95.4% and PPV of 83.8% on the subset of 11,990 hospitalized patients with discharge diagnosis codes. We successfully add nearly 20,000 labels to CTPAs in a publicly available dataset and demonstrate the external validity of a semi-supervised language model in accelerating hematologic research.

肺栓塞自然语言处理医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。