用入院记录预测中风后功能恢复,大模型表现媲美传统方法。
Large Language Models Predict Functional Outcomes after Acute Ischemic Stroke
- 用大语言模型直接分析入院记录预测中风结局。
- 90天准确率33.9%(精确分类),76.3%(功能恢复二分类)。
- 无需手动提取数据,可嵌入临床流程,适合医疗AI研发者。
急性缺血性中风后功能预后的准确预测有助于临床决策与资源分配。以往研究多依赖年龄、NIHSS等结构化变量和传统机器学习模型。大语言模型(LLMs)能否直接从常规入院记录中推断未来改良秩次量表(mRS)评分,仍缺乏探索。本研究评估了编码器(BERT、NYUTron)与生成式模型(Llama-3.1-8B、MedGemma-4B)在冻结与微调设置下的表现,使用纽约大学朗格尼中风注册库(2016–2025)数据,按时间划分,最近12个月用于测试。出院结局数据集含9,485份病史体格检查记录,90天结局数据集含1,898份记录。以7类精确mRS准确率及二分类功能结局(mRS 0–2 vs. 3–6)准确率为评估指标,对比包含NIHSS与年龄的结构化基线模型。微调后的Llama在90天预测中表现最佳,精确准确率达33.9% [95% CI, 27.9–39.9%],二分类准确率为76.3% [95% CI, 70.7–81.9%];出院预测精确准确率为42.0% [95% CI, 39.0–45.0%],二分类准确率为75.0% [95% CI, 72.4–77.6%]。90天预测性能与结构化基线相当。结果表明,微调后的大型语言模型仅凭入院记录即可实现与需结构化变量抽象模型相当的中风后功能预后预测能力,支持开发无需人工数据提取、可无缝集成于临床流程的文本驱动预后工具。
原文摘要 · Abstract (English)
Accurate prediction of functional outcomes after acute ischemic stroke can inform clinical decision-making and resource allocation. Prior work on modified Rankin Scale (mRS) prediction has relied primarily on structured variables (e.g., age, NIHSS) and conventional machine learning. The ability of large language models (LLMs) to infer future mRS scores directly from routine admission notes remains largely unexplored. We evaluated encoder (BERT, NYUTron) and generative (Llama-3.1-8B, MedGemma-4B) LLMs, in both frozen and fine-tuned settings, for discharge and 90-day mRS prediction using a large, real-world stroke registry. The discharge outcome dataset included 9,485 History and Physical notes and the 90-day outcome dataset included 1,898 notes from the NYU Langone Get With The Guidelines-Stroke registry (2016-2025). Data were temporally split with the most recent 12 months held out for testing. Performance was assessed using exact (7-class) mRS accuracy and binary functional outcome (mRS 0-2 vs. 3-6) accuracy and compared against established structured-data baselines incorporating NIHSS and age. Fine-tuned Llama achieved the highest performance, with 90-day exact mRS accuracy of 33.9% [95% CI, 27.9-39.9%] and binary accuracy of 76.3% [95% CI, 70.7-81.9%]. Discharge performance reached 42.0% [95% CI, 39.0-45.0%] exact accuracy and 75.0% [95% CI, 72.4-77.6%] binary accuracy. For 90-day prediction, Llama performed comparably to structured-data baselines. Fine-tuned LLMs can predict post-stroke functional outcomes from admission notes alone, achieving performance comparable to models requiring structured variable abstraction. Our findings support the development of text-based prognostic tools that integrate seamlessly into clinical workflows without manual data extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。