arXiv:2510.16536q-bio.QMcs.AI2025-10

用少量标注数据融合基因与心电信息,用大模型预测心血管风险。

Few-Label Multimodal Modeling of SNP Variants and ECG Phenotypes Using Large Language Models for Cardiovascular Risk Stratification

  • 借大模型融合基因型与心电图特征,弱化标注依赖。
  • 仅需少量真实标签即达全量数据训练模型性能。
  • 生成临床可解释推理链,适合医疗决策支持场景。

心血管疾病(CVD)风险分层因多因素性及高质量标注数据稀缺而面临挑战。尽管基因组和心电生理数据(如SNP变异和心电图表型)日益可得,但在低标注条件下有效整合多模态信息仍具难度。这源于高质量多模态数据集稀缺及生物信号高维性,制约传统监督模型效果。为此,我们提出一种少标签多模态框架,利用大语言模型(LLM)融合遗传与电生理信息进行心血管风险分层。方法引入伪标签精炼策略,自适应地从弱监督预测中提取高置信度标签,实现仅需少量真实标注即可稳健微调模型。为提升可解释性,将任务设为思维链(CoT)推理问题,使模型输出临床相关推理解释。实验表明,多模态输入、少标签监督与CoT推理的结合显著提升模型在多样患者群体中的鲁棒性与泛化能力。使用多模态SNP变异与心电图衍生特征的实验结果表明,该方法性能可媲美全量数据训练模型,凸显基于大模型的少标签多模态建模在推进个性化心血管诊疗中的潜力。

原文摘要 · Abstract (English)

Cardiovascular disease (CVD) risk stratification remains a major challenge due to its multifactorial nature and limited availability of high-quality labeled datasets. While genomic and electrophysiological data such as SNP variants and ECG phenotypes are increasingly accessible, effectively integrating these modalities in low-label settings is non-trivial. This challenge arises from the scarcity of well-annotated multimodal datasets and the high dimensionality of biological signals, which limit the effectiveness of conventional supervised models. To address this, we present a few-label multimodal framework that leverages large language models (LLMs) to combine genetic and electrophysiological information for cardiovascular risk stratification. Our approach incorporates a pseudo-label refinement strategy to adaptively distill high-confidence labels from weakly supervised predictions, enabling robust model fine-tuning with only a small set of ground-truth annotations. To enhance the interpretability, we frame the task as a Chain of Thought (CoT) reasoning problem, prompting the model to produce clinically relevant rationales alongside predictions. Experimental results demonstrate that the integration of multimodal inputs, few-label supervision, and CoT reasoning improves robustness and generalizability across diverse patient profiles. Experimental results using multimodal SNP variants and ECG-derived features demonstrated comparable performance to models trained on the full dataset, underscoring the promise of LLM-based few-label multimodal modeling for advancing personalized cardiovascular care.

心血管风险多模态融合小样本学习大模型医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。