用出院记录精准分类心衰患者类型,且模型解释能对得上医生判断。
Interpretable phenotyping of Heart Failure patients with Dutch discharge letters
- 用大模型分析出院记录文本,自动划分心衰患者类型。
- 仅用文本数据就达到AUC 0.84的分类效果,超越传统方法。
- 新模型解释结果与医生判断高度一致,适合临床可信应用。
心衰患者表现多样,影响治疗和预后。本研究基于左心室射血分数(LVEF)类别,评估了利用结构化与非结构化数据进行心衰患者表型分型的模型性能与可解释性。研究纳入2015至2023年阿姆斯特丹大学医学中心(AMC和VUmc)共33,105次心衰住院记录(16,334名患者),以AMC数据训练,VUmc数据外部验证。数据包括临床指标与出院记录,银标准标签通过诊断编码、超声心动图结果及文本提及综合生成,金标准由两名医生手动标注300例患者。训练并比较了基于Transformer的黑箱模型与增广线性(Aug-Linear)白箱模型,以及基线模型。为评估可解释性,两名临床医生标注20份出院记录中用于判断LVEF的关键信息,并与黑箱模型的SHAP/LIME解释及Aug-Linear模型的内在解释对比。结果显示,仅使用出院记录的BERT与Aug-Linear模型在外部验证中表现最佳(AUC分别为0.84与0.81),优于基线;Aug-Linear的解释与医生标注更接近。结论:出院记录是心衰表型分型最有效信息源,Aug-Linear模型在保持黑箱性能的同时提供与医生一致的解释,支持其在透明临床决策中的应用。
原文摘要 · Abstract (English)
Objective: Heart failure (HF) patients present with diverse phenotypes affecting treatment and prognosis. This study evaluates models for phenotyping HF patients based on left ventricular ejection fraction (LVEF) classes, using structured and unstructured data, assessing performance and interpretability. Materials and Methods: The study analyzes all HF hospitalizations at both Amsterdam UMC hospitals (AMC and VUmc) from 2015 to 2023 (33,105 hospitalizations, 16,334 patients). Data from AMC were used for model training, and from VUmc for external validation. The dataset was unlabelled and included tabular clinical measurements and discharge letters. Silver labels for LVEF classes were generated by combining diagnosis codes, echocardiography results, and textual mentions. Gold labels were manually annotated for 300 patients for testing. Multiple Transformer-based (black-box) and Aug-Linear (white-box) models were trained and compared with baselines on structured and unstructured data. To evaluate interpretability, two clinicians annotated 20 discharge letters by highlighting information they considered relevant for LVEF classification. These were compared to SHAP and LIME explanations from black-box models and the inherent explanations of Aug-Linear models. Results: BERT-based and Aug-Linear models, using discharge letters alone, achieved the highest classification results (AUC=0.84 for BERT, 0.81 for Aug-Linear on external validation), outperforming baselines. Aug-Linear explanations aligned more closely with clinicians' explanations than post-hoc explanations on black-box models. Conclusions: Discharge letters emerged as the most informative source for phenotyping HF patients. Aug-Linear models matched black-box performance while providing clinician-aligned interpretability, supporting their use in transparent clinical decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。