arXiv:2505.23879cs.LGcs.AI2025-05被引 3

用刺突蛋白序列和临床数据预测新冠重症,准确率超82%

CNN-LSTM Hybrid Model for AI-Driven Prediction of COVID-19 Severity from Spike Sequences and Clinical Data

  • 结合卷积与循环神经网络,捕捉基因序列局部特征与时间依赖性
  • 在9570条序列上实现82.92%的F1分数,性能稳定无过拟合
  • 适用于疫情监测与精准医疗,为未来突发病提供预警框架

新冠疫情由严重急性呼吸综合征冠状病毒2(SARS-CoV-2)引起,准确预测疾病严重程度对优化医疗资源配置和患者管理至关重要。刺突蛋白介导病毒入侵宿主细胞,其受体结合域突变率高,影响病毒致病性。人工智能方法如深度学习可整合基因组与临床数据以预测疾病结局。本研究旨在构建一个混合CNN-LSTM深度学习模型,基于南美患者的数据,利用刺突蛋白序列与临床元数据预测新冠严重程度。从GISAID数据库获取9,570条刺突蛋白序列,经标准化后筛选出3,467条符合标准的序列,包含2,313例重症与1,154例轻症。通过特征工程提取序列特征,对人口统计学与临床变量进行独热编码。采用混合CNN-LSTM架构,结合卷积层提取局部模式与长短期记忆层建模长期依赖关系。模型在测试集上达到F1分数82.92%、ROC-AUC 0.9084、精确率83.56%、召回率82.85%,表现稳健且无明显过拟合。最常见谱系(P.1、AY.99.2)与分支(GR、GK)与区域流行趋势一致,提示病毒遗传特征可能与临床结果相关。结论表明,该混合模型能有效利用刺突蛋白序列与临床数据预测新冠严重程度,凸显人工智能在基因组监测与精准公共卫生中的价值。尽管存在局限性,该方法为未来疫情早期严重程度预测提供了可行框架。

原文摘要 · Abstract (English)

The COVID-19 pandemic, caused by SARS-CoV-2, highlighted the critical need for accurate prediction of disease severity to optimize healthcare resource allocation and patient management. The spike protein, which facilitates viral entry into host cells, exhibits high mutation rates, particularly in the receptor-binding domain, influencing viral pathogenicity. Artificial intelligence approaches, such as deep learning, offer promising solutions for leveraging genomic and clinical data to predict disease outcomes. Objective: This study aimed to develop a hybrid CNN-LSTM deep learning model to predict COVID-19 severity using spike protein sequences and associated clinical metadata from South American patients. Methods: We retrieved 9,570 spike protein sequences from the GISAID database, of which 3,467 met inclusion criteria after standardization. The dataset included 2,313 severe and 1,154 mild cases. A feature engineering pipeline extracted features from sequences, while demographic and clinical variables were one-hot encoded. A hybrid CNN-LSTM architecture was trained, combining CNN layers for local pattern extraction and an LSTM layer for long-term dependency modeling. Results: The model achieved an F1 score of 82.92%, ROC-AUC of 0.9084, precision of 83.56%, and recall of 82.85%, demonstrating robust classification performance. Training stabilized at 85% accuracy with minimal overfitting. The most prevalent lineages (P.1, AY.99.2) and clades (GR, GK) aligned with regional epidemiological trends, suggesting potential associations between viral genetics and clinical outcomes. Conclusion: The CNN-LSTM hybrid model effectively predicted COVID-19 severity using spike protein sequences and clinical data, highlighting the utility of AI in genomic surveillance and precision public health. Despite limitations, this approach provides a framework for early severity prediction in future outbreaks.

新冠预测深度学习基因序列临床数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。