用大模型实现可自适应的心血管病风险预测,支持多种数据类型和人群。
Adaptable Cardiovascular Disease Risk Prediction from Heterogeneous Data using Large Language Models
- 基于大语言模型,灵活整合结构化与非结构化医疗数据。
- 在超50万英国人数据上训练,表现优于传统模型和现有风险评分。
- 仅需少量新数据即可快速适配新人群,适合真实复杂医疗场景。
心血管疾病(CVD)风险预测模型对识别高危个体和指导预防措施至关重要。然而,现有模型在真实临床实践中面临挑战:患者画像过于简化、输入格式僵化,且对分布偏移敏感。我们开发了AdaCVD——一个基于大规模语言模型、在英国生物样本库超过50万参与者数据上充分微调的可自适应CVD风险预测框架。基准对比显示,AdaCVD超越了现有风险评分和标准机器学习方法,达到领先水平。关键突破在于首次在三个维度解决核心临床难题:灵活整合全面但异质的患者信息;无缝融合结构化数据与非结构化文本;仅用少量新增数据即可快速适应新患者群体。分层分析表明,其在不同人口统计、社会经济及临床亚组中均表现稳健,包括代表性不足的群体。AdaCVD为构建更灵活、智能化的临床决策支持工具提供了可行路径,适用于异质性与动态性强的医疗环境。
原文摘要 · Abstract (English)
Cardiovascular disease (CVD) risk prediction models are essential for identifying high-risk individuals and guiding preventive actions. However, existing models struggle with the challenges of real-world clinical practice as they oversimplify patient profiles, rely on rigid input schemas, and are sensitive to distribution shifts. We developed AdaCVD, an adaptable CVD risk prediction framework built on large language models extensively fine-tuned on over half a million participants from the UK Biobank. In benchmark comparisons, AdaCVD surpasses established risk scores and standard machine learning approaches, achieving state-of-the-art performance. Crucially, for the first time, it addresses key clinical challenges across three dimensions: it flexibly incorporates comprehensive yet variable patient information; it seamlessly integrates both structured data and unstructured text; and it rapidly adapts to new patient populations using minimal additional data. In stratified analyses, it demonstrates robust performance across demographic, socioeconomic, and clinical subgroups, including underrepresented cohorts. AdaCVD offers a promising path toward more flexible, AI-driven clinical decision support tools suited to the realities of heterogeneous and dynamic healthcare environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。