用大模型分析心脏病死亡预测的代价收益,让医生看得懂、选得准。
Cost-Aware Prediction (CAP): An LLM-Enhanced Machine Learning Pipeline and Decision Support System for Heart Failure Mortality Prediction
- 结合大模型生成患者个体化成本效益分析,辅助临床决策。
- 模型在3万例心衰患者上实现0.804的AUROC和0.135的Brier分数。
- 适合关注医疗决策透明性与可解释性的临床研究者使用。
目标:机器学习预测模型常忽视下游价值权衡与临床可解释性。本文提出一种成本感知预测(CAP)框架,利用大语言模型(LLM)代理辅助进行成本-收益分析,以沟通应用机器学习预测时的权衡。方法:基于30,021名心衰患者数据(22%一年内死亡率),构建预测模型以识别适合居家护理的患者;引入临床影响投影(CIP)曲线,可视化生活质量及医疗成本(治疗与错误成本)等关键维度,评估预测的临床后果;最后使用四个LLM代理生成个性化患者描述。系统经临床医生评估其决策支持价值。结果:梯度提升树(XGB)模型表现最佳,受试者工作特征曲线下面积(AUROC)为0.804(95%置信区间0.792–0.816),精确率-召回率曲线下面积(AUPRC)为0.529(95%置信区间0.502–0.558),布莱尔得分(Brier score)为0.135(95%置信区间0.130–0.140)。讨论:CIP曲线提供不同决策阈值下的群体成本构成概览,而LLM在个体层面生成成本-收益分析。临床医生评价积极,但反馈指出需增强对推测性任务的技术准确性。结论:CAP通过整合机器学习输出与成本-收益分析,提升决策支持的透明性与可解释性。
原文摘要 · Abstract (English)
Objective: Machine learning (ML) predictive models are often developed without considering downstream value trade-offs and clinical interpretability. This paper introduces a cost-aware prediction (CAP) framework that combines cost-benefit analysis assisted by large language model (LLM) agents to communicate the trade-offs involved in applying ML predictions. Materials and Methods: We developed an ML model predicting 1-year mortality in patients with heart failure (N = 30,021, 22% mortality) to identify those eligible for home care. We then introduced clinical impact projection (CIP) curves to visualize important cost dimensions - quality of life and healthcare provider expenses, further divided into treatment and error costs, to assess the clinical consequences of predictions. Finally, we used four LLM agents to generate patient-specific descriptions. The system was evaluated by clinicians for its decision support value. Results: The eXtreme gradient boosting (XGB) model achieved the best performance, with an area under the receiver operating characteristic curve (AUROC) of 0.804 (95% confidence interval (CI) 0.792-0.816), area under the precision-recall curve (AUPRC) of 0.529 (95% CI 0.502-0.558) and a Brier score of 0.135 (95% CI 0.130-0.140). Discussion: The CIP cost curves provided a population-level overview of cost composition across decision thresholds, whereas LLM-generated cost-benefit analysis at individual patient-levels. The system was well received according to the evaluation by clinicians. However, feedback emphasizes the need to strengthen the technical accuracy for speculative tasks. Conclusion: CAP utilizes LLM agents to integrate ML classifier outcomes and cost-benefit analysis for more transparent and interpretable decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。