大语言模型能提升医疗预测,但需解决公平性与可解释性难题。
Will Large Language Models Transform Clinical Prediction?
- 利用大语言模型处理长期电子病历数据,支持多疾病多结局预测。
- 现有模型存在校准差、外部验证不足及对少数群体的偏见问题。
- 适合关注医疗AI落地、临床决策系统优化的研究者阅读。
目的:大语言模型(LLMs)在医疗领域日益受到关注。本文评述了其在改进诊断与预后任务的临床预测模型(CPMs)方面的潜力,尤其聚焦于处理纵向电子健康记录(EHR)数据的能力。发现:LLMs在处理多模态和纵向EHR数据方面展现出前景,可支持多种健康状况的多结局预测。然而,方法学、验证、基础设施和监管方面仍存挑战,包括时间-事件建模方法不完善、预测校准度差、外部验证有限,以及对代表性不足群体的偏见。高昂的基础设施成本和缺乏明确的监管框架也阻碍了实际应用。启示:需进一步研究与跨学科合作,以实现公平且有效的临床预测集成。开发具备时间感知、公平性和可解释性的模型应成为转型临床预测流程的重点。
原文摘要 · Abstract (English)
Objective: Large language models (LLMs) are attracting increasing interest in healthcare. This commentary evaluates the potential of LLMs to improve clinical prediction models (CPMs) for diagnostic and prognostic tasks, with a focus on their ability to process longitudinal electronic health record (EHR) data. Findings: LLMs show promise in handling multimodal and longitudinal EHR data and can support multi-outcome predictions for diverse health conditions. However, methodological, validation, infrastructural, and regulatory chal- lenges remain. These include inadequate methods for time-to-event modelling, poor calibration of predictions, limited external validation, and bias affecting underrepresented groups. High infrastructure costs and the absence of clear regulatory frameworks further prevent adoption. Implications: Further work and interdisciplinary collaboration are needed to support equitable and effective integra- tion into the clinical prediction. Developing temporally aware, fair, and explainable models should be a priority focus for transforming clinical prediction workflow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。