用大模型连接基因变异与心电图特征,提升心血管风险预测
How Effectively Can Large Language Models Connect SNP Variants and ECG Phenotypes for Cardiovascular Risk Prediction?
- 将基因数据转为思维链任务,让大模型推理疾病关联
- 能从家族遗传模式中挖掘潜在致病基因突变
- 适合临床科研与个性化医疗研究者参考
心血管疾病(CVD)预测因多因素病因和全球高发病率、死亡率而极具挑战。尽管基因组和心电图数据日益丰富,但从高维、噪声大且标注稀疏的数据中提取生物学意义仍十分困难。本文探索微调后的大语言模型(LLM)在利用高通量基因组分析获得的遗传标记预测心脏病及潜在导致CVD风险的单核苷酸多态性(SNP)方面的潜力。通过将问题建模为思维链(Chain of Thought, CoT)推理任务,模型被提示生成疾病标签并针对不同患者表型做出有依据的临床推断。研究发现,模型能够从结构化与半结构化的基因数据中学习隐含的生物关系,尤其在基于家族遗传模式的分析中表现突出。结果表明,大模型在早期检测、风险评估及推动个性化心脏诊疗方面具有重要前景。
原文摘要 · Abstract (English)
Cardiovascular disease (CVD) prediction remains a tremendous challenge due to its multifactorial etiology and global burden of morbidity and mortality. Despite the growing availability of genomic and electrophysiological data, extracting biologically meaningful insights from such high-dimensional, noisy, and sparsely annotated datasets remains a non-trivial task. Recently, LLMs has been applied effectively to predict structural variations in biological sequences. In this work, we explore the potential of fine-tuned LLMs to predict cardiac diseases and SNPs potentially leading to CVD risk using genetic markers derived from high-throughput genomic profiling. We investigate the effect of genetic patterns associated with cardiac conditions and evaluate how LLMs can learn latent biological relationships from structured and semi-structured genomic data obtained by mapping genetic aspects that are inherited from the family tree. By framing the problem as a Chain of Thought (CoT) reasoning task, the models are prompted to generate disease labels and articulate informed clinical deductions across diverse patient profiles and phenotypes. The findings highlight the promise of LLMs in contributing to early detection, risk assessment, and ultimately, the advancement of personalized medicine in cardiac care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。