用大模型生成的专家摘要增强重症患者死亡预测,效果显著但多为信息重组。
Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries
- 融合生理数据、临床笔记与大模型生成摘要进行死亡预测
- 联合模型AUPRC达0.4977,较仅用生理数据提升20%
- 摘要主要重排已有信息,非引入新内容,适合医疗AI研究者
为评估一种多表征框架在重症监护室(ICU)患者住院死亡率(IHM)预测中的表现,将大语言模型(LLM)生成的专家摘要与生理数据融合,并分析其增益是否与原始病历冗余。基于MIMIC-III数据集(19,211例首次入住ICU患者,死亡率12.83%),对48小时生理数据、临床记录及在禁止预判提示下生成的LLM摘要进行编码并融合。通过岭回归可恢复性、线性探测和替换消融实验(以病历正交残差或患者随机摘要嵌入替代)评估冗余性。在3,843个保留病例上,融合模型达到AUPRC 0.4977/AUROC 0.8429,优于仅用生理数据的0.3625/0.7770。从病历嵌入中可解释摘要嵌入方差的40.8%。病历正交残差仍保留部分增益(+0.0258 AUPRC,95%置信区间0.005–0.047;占比28%),而患者随机摘要则低于基准。结果表明,摘要能提升个体化预测,但主要作用是重组病历中已有的信息。
原文摘要 · Abstract (English)
To evaluate a multi-representational framework in which large language model (LLM)-generated expert summaries of intensive care unit (ICU) notes are fused with physiology for in-hospital mortality (IHM) prediction, and to determine how much of the resulting gain is non-redundant with the notes themselves. Using MIMIC-III (19,211 first ICU stays, 12.83% mortality), we encoded 48-hour physiology, clinical notes, and LLM summaries generated under a prompt forbidding prognostication, then fused them. Redundancy was assessed by ridge recoverability, linear probes, and retrained ablations substituting a note-orthogonal residual or a patient-shuffled summary embedding. On 3,843 held-out stays, fusion reached AUPRC 0.4977/AUROC 0.8429 versus 0.3625/0.7770 for physiology alone. Ridge regression from note embeddings explained 40.8% of summary-embedding variance. The note-orthogonal residual retained a minority of the gain (+0.0258 AUPRC, 95% CI 0.005--0.047; 28%), while patient-shuffled summaries fell below the reference. Summaries improve prediction patient-specifically, but predominantly by reorganizing information the notes already contain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。