用混合NLP技术从病历中高效提取新冠后遗症症状,提升诊断准确率。
Extracting Post-Acute Sequelae of SARS-CoV-2 Infection Symptoms from Clinical Notes via Hybrid Natural Language Processing
- 结合规则与BERT模型,识别病历中的新冠后遗症症状及其存在状态。
- 内部验证F1达0.82,跨机构外部验证F1为0.76,结果稳定可靠。
- 适合临床研究者和医疗数据工程师用于大规模后遗症筛查。
由于新冠后遗症(PASC)症状多样且持续时间不一,准确高效诊断仍具挑战。为此,我们构建了一个融合规则驱动的命名实体识别与BERT-based断言检测模块的混合自然语言处理流程,用于从临床笔记中提取并判断PASC症状的存在性。我们联合临床专家制定了全面的PASC术语词典。基于美国RECOVER计划网络中11个医疗系统的数据,我们整理了160份入院记录用于模型开发与评估,并收集了47,654份进度记录开展人群水平患病率研究。在单机构内部验证中,断言检测平均F1得分为0.82;在10个外部机构的交叉验证中为0.76。模型平均每篇笔记处理耗时2.448±0.812秒。斯皮尔曼相关性分析显示,阳性提及相关系数ρ>0.83,阴性提及ρ>0.72,均具有统计学显著性(P<0.0001)。结果表明该模型在准确性和效率上均表现优异,具备改善PASC诊断的潜力。
原文摘要 · Abstract (English)
Accurately and efficiently diagnosing Post-Acute Sequelae of COVID-19 (PASC) remains challenging due to its myriad symptoms that evolve over long- and variable-time intervals. To address this issue, we developed a hybrid natural language processing pipeline that integrates rule-based named entity recognition with BERT-based assertion detection modules for PASC-symptom extraction and assertion detection from clinical notes. We developed a comprehensive PASC lexicon with clinical specialists. From 11 health systems of the RECOVER initiative network across the U.S., we curated 160 intake progress notes for model development and evaluation, and collected 47,654 progress notes for a population-level prevalence study. We achieved an average F1 score of 0.82 in one-site internal validation and 0.76 in 10-site external validation for assertion detection. Our pipeline processed each note at $2.448\pm 0.812$ seconds on average. Spearman correlation tests showed $ρ>0.83$ for positive mentions and $ρ>0.72$ for negative ones, both with $P <0.0001$. These demonstrate the effectiveness and efficiency of our models and their potential for improving PASC diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。