用大模型嵌入分析临床记录,提前预测创伤后癫痫风险。
Predicting Post-Traumatic Epilepsy from Clinical Records using Large Language Model Embeddings

- 用大语言模型提取临床记录的语义特征,融合表格数据进行预测。
- 最佳模型AUC达0.892,关键因素包括急性发作、伤情严重度等。
- 无需昂贵影像,适合临床早期筛查,尤其对资源有限地区有帮助。
目的:创伤后癫痫(PTE)是脑外伤(TBI)后一种致残性神经疾病,早期预测因临床数据异质性强、阳性病例少及依赖高成本神经影像而困难。本文研究仅使用常规急性期临床记录,结合基于大语言模型的方法是否可实现早期PTE预测。方法:基于TRACK-TBI队列的筛选数据集,构建自动化预测框架,采用预训练大语言模型(LLM)作为固定特征提取器,编码临床记录文本。对比纯表格特征、LLM嵌入及混合特征表示在分层交叉验证下梯度提升树分类器的表现。结果:与仅使用表格特征相比,LLM嵌入能更有效捕捉临床上下文信息。模态感知特征融合策略表现最优,达到AUC-ROC 0.892,AUPRC 0.798。急性创伤后癫痫发作、伤情严重程度、神经外科干预和重症监护室住院时间是主要贡献因素。意义:研究表明,常规急性期临床记录中蕴含可用于早期PTE风险预测的信息,结合LLM嵌入与梯度提升树分类器的方法具有潜力,可作为影像基预测的有力补充。
原文摘要 · Abstract (English)
Objective: Post-traumatic epilepsy (PTE) is a debilitating neurological disorder that develops after traumatic brain injury (TBI). Early prediction of PTE remains challenging due to heterogeneous clinical data, limited positive cases, and reliance on resource-intensive neuroimaging data. We investigate whether routinely collected acute clinical records alone can support early PTE prediction using language model-based approaches. Methods: Using a curated subset of the TRACK-TBI cohort, we developed an automated PTE prediction framework that implements pretrained large language models (LLMs) as fixed feature extractors to encode clinical records. Tabular features, LLM-generated embeddings, and hybrid feature representations were evaluated using gradient-boosted tree classifiers under stratified cross-validation. Results: LLM embeddings achieved performance improvements by capturing contextual clinical information compared to using tabular features alone. The best performance was achieved by a modality-aware feature fusion strategy combining tabular features and LLM embeddings, achieving an AUC-ROC of 0.892 and AUPRC of 0.798. Acute post-traumatic seizures, injury severity, neurosurgical intervention, and ICU stay are key contributors to the predictive performance. Significance: These findings demonstrate that routine acute clinical records contain information suitable for early PTE risk prediction using LLM embeddings in conjunction with gradient-boosted tree classifiers. This approach represents a promising complement to imaging-based prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。