用可解释的神经网络发现新冠突变与毒力关联,比传统方法更准。
Explainable convolutional neural network model provides an alternative genome-wide association perspective on mutations in SARS-CoV-2
- 用可解释的卷积神经网络预测变异株,识别关键突变
- 在刺突基因区发现已知重要突变,效果优于传统分析
- 适合研究病毒进化与突变影响的生物信息学工作者
识别与SARS-CoV-2毒株表型变化相关的突变,对疫情预测和防控至关重要。我们比较了可解释卷积神经网络(CNN)方法与传统的全基因组关联研究(GWAS),分析与世卫组织(WHO)分类的变异株(作为毒力表型的代理)相关的突变。我们训练了一个CNN分类模型,能够根据基因组序列预测变异株类型,并使用SHAP(Shapley Additive Explanations)模型识别对正确预测至关重要的突变。作为对比,我们还进行了传统的GWAS分析以识别与变异株相关的突变。结果表明,可解释的神经网络方法能更有效地揭示已知与变异株相关的核苷酸替换,例如刺突基因区域的突变。这些结果表明,用于基因组序列的可解释神经网络为传统的全基因组分析提供了有前景的替代方案。
原文摘要 · Abstract (English)
Identifying mutations of SARS-CoV-2 strains associated with their phenotypic changes is critical for pandemic prediction and prevention. We compared an explainable convolutional neural network (CNN) approach and the traditional genome-wide association study (GWAS) on the mutations associated with WHO labels of SARS-CoV-2, a proxy for virulence phenotypes. We trained a CNN classification model that can predict genomic sequences into Variants of Concern (VOCs) and then applied Shapley Additive explanations (SHAP) model to identify mutations that are important for the correct predictions. For comparison, we performed traditional GWAS to identify mutations associated with VOCs. Comparison of the two approaches shows that the explainable neural network approach can more effectively reveal known nucleotide substitutions associated with VOCs, such as those in the spike gene regions. Our results suggest that explainable neural networks for genomic sequences offer a promising alternative to the traditional genome wide analysis approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。