用混合框架大幅降低算力需求,高效识别痴呆患者。
Accelerating Clinical NLP at Scale with a Hybrid Framework with Reduced GPU Demands: A Case Study in Dementia Identification
- 结合规则过滤、SVM与BERT的混合模型,兼顾效率与精度。
- 在490万退伍军人数据中实现F1=0.87,检出数超结构化方法3倍。
- 单机双A40 GPU两周完成,适合资源有限的医疗机构。
临床自然语言处理(NLP)在临床研究和实践中日益重要,但多数先进方案基于Transformer,需高算力支持,限制了普及。本文提出一种混合NLP框架,融合规则过滤、支持向量机(SVM)分类器与BERT模型,在保持准确率的同时提升效率。该框架应用于痴呆识别案例,分析490万患新发高血压的退伍军人,涵盖21亿份临床记录。在患者层面,方法达到0.90的精确率、0.84的召回率和0.87的F1分数。此外,该NLP方法检出的痴呆病例数超过结构化数据方法的三倍。所有处理在配备双A40 GPU的单机上约两周内完成。本研究证明了混合NLP方案在大规模临床文本分析中的可行性,使前沿技术更易为算力有限的医疗单位所用。
原文摘要 · Abstract (English)
Clinical natural language processing (NLP) is increasingly in demand in both clinical research and operational practice. However, most of the state-of-the-art solutions are transformers-based and require high computational resources, limiting their accessibility. We propose a hybrid NLP framework that integrates rule-based filtering, a Support Vector Machine (SVM) classifier, and a BERT-based model to improve efficiency while maintaining accuracy. We applied this framework in a dementia identification case study involving 4.9 million veterans with incident hypertension, analyzing 2.1 billion clinical notes. At the patient level, our method achieved a precision of 0.90, a recall of 0.84, and an F1-score of 0.87. Additionally, this NLP approach identified over three times as many dementia cases as structured data methods. All processing was completed in approximately two weeks using a single machine with dual A40 GPUs. This study demonstrates the feasibility of hybrid NLP solutions for large-scale clinical text analysis, making state-of-the-art methods more accessible to healthcare organizations with limited computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。