用大模型自动识别放疗安全事件报告,跨机构效果显著提升。
Automated Triaging and Transfer Learning of Incident Learning Safety Reports Using Large Language Representational Models
- 用蓝伯特大模型+迁移学习,从两院数据中训练智能筛选系统。
- 跨机构测试下模型准确率从0.56提升至0.78,接近人工水平。
- 适合医疗质控团队快速筛查高危事件,减少人工负担。
目的:事故报告是医疗安全与质量改进的重要工具,但人工审阅耗时且需专业经验。本文提出一种自然语言处理(NLP)筛查工具,用于在两家机构的放射治疗领域检测高严重性事故报告。方法与材料:使用两个文本数据集训练和评估模型:本机构7,094份报告(Inst.),以及国际原子能机构SAFRON的571份报告(SF),所有报告均经临床专家标注严重性评分。训练并评估两类模型:基线支持向量机(SVM)和预训练于医学文献及住院患者数据的BlueBERT大语言模型。通过两种方式评估模型泛化能力:一是在Inst.-train上训练,于SF-test上测试;二是训练一个先在Inst.-train微调、再在SF-train微调的BlueBERT_TRANSFER模型,最后在SF-test集上测试。为进一步分析性能,还对本机构59份经人工优化清晰度的报告进行测试。结果:在Inst.测试集上,SVM和BlueBERT的分类表现分别为AUROC 0.82和0.81。未采用跨机构迁移学习时,于SF测试集上的表现受限于SVM的AUROC 0.42和BlueBERT的0.56。而经过双阶段微调的BlueBERT_TRANSFER模型在SF测试集上提升至AUROC 0.78。在手动优化的本机构报告上,SVM与BlueBERT_TRANSFER的性能分别为AUROC 0.85和0.74,与人类专家表现(AUROC 0.81)相当。结论:我们成功开发了基于事故报告文本的跨机构放射治疗领域NLP模型,可在优化数据集上达到与人工相似的高严重性事件识别能力。
原文摘要 · Abstract (English)
PURPOSE: Incident reports are an important tool for safety and quality improvement in healthcare, but manual review is time-consuming and requires subject matter expertise. Here we present a natural language processing (NLP) screening tool to detect high-severity incident reports in radiation oncology across two institutions. METHODS AND MATERIALS: We used two text datasets to train and evaluate our NLP models: 7,094 reports from our institution (Inst.), and 571 from IAEA SAFRON (SF), all of which had severity scores labeled by clinical content experts. We trained and evaluated two types of models: baseline support vector machines (SVM) and BlueBERT which is a large language model pretrained on PubMed abstracts and hospitalized patient data. We assessed for generalizability of our model in two ways. First, we evaluated models trained using Inst.-train on SF-test. Second, we trained a BlueBERT_TRANSFER model that was first fine-tuned on Inst.-train then on SF-train before testing on SF-test set. To further analyze model performance, we also examined a subset of 59 reports from our Inst. dataset, which were manually edited for clarity. RESULTS Classification performance on the Inst. test achieved AUROC 0.82 using SVM and 0.81 using BlueBERT. Without cross-institution transfer learning, performance on the SF test was limited to an AUROC of 0.42 using SVM and 0.56 using BlueBERT. BlueBERT_TRANSFER, which was fine-tuned on both datasets, improved the performance on SF test to AUROC 0.78. Performance of SVM, and BlueBERT_TRANSFER models on the manually curated Inst. reports (AUROC 0.85 and 0.74) was similar to human performance (AUROC 0.81). CONCLUSION: In summary, we successfully developed cross-institution NLP models on incident report text from radiation oncology centers. These models were able to detect high-severity reports similarly to humans on a curated dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。