用生物先验知识提升基因组预测,让模型在少数据下更准且可解释。
Biology-informed neural networks learn nonlinear representations from omics data to improve genomic prediction and interpretability
- 将基因型与多组学数据结合,训练时用生物通路先验,推理时仅需基因型。
- 在稀疏数据下,预测相关性提升56%,比传统方法更准确。
- 能发现全基因关联研究忽略的非线性关键基因和通路,适合育种研究。
我们扩展了生物信息神经网络(BINNs)用于作物基因组预测(GP)和选择(GS),将数千个单核苷酸多态性(SNPs)与多组学测量及先验生物学知识融合。传统基因型-表型(G2P)模型依赖直接映射,准确率有限,迫使育种者进行大规模、高成本田间试验以维持或小幅提升遗传增益。包含基因表达等中间分子表型的模型虽能提高预测拟合度,但因部署时无法获取此类数据而难以用于基因组选择。BINNs通过编码通路级归纳偏置,在训练时利用多组学数据,推理时仅使用基因型数据克服此限制。应用于玉米基因表达和多环境田间试验数据,BINN在子群体内及跨子群体的稀疏数据条件下,排名相关性准确率最高提升56%,并非线性识别出GWAS/TWAS未能发现的基因。在合成代谢组学基准中,具备完整领域知识时,相比传统神经网络,预测误差降低75%,且正确识别出最重要的非线性通路。重要的是,两个案例中高度敏感的隐变量均与所代表的实验量高度相关,即使未直接训练。这表明BINNs能从基因型到表型学习具有生物学意义的表示,无论线性或非线性。综上,BINNs建立了一个框架,利用中间领域信息提升基因组预测精度,并揭示可指导基因组选择、候选基因筛选、通路富集分析和基因编辑优先级的非线性生物关系。
原文摘要 · Abstract (English)
We extend biologically-informed neural networks (BINNs) for genomic prediction (GP) and selection (GS) in crops by integrating thousands of single-nucleotide polymorphisms (SNPs) with multi-omics measurements and prior biological knowledge. Traditional genotype-to-phenotype (G2P) models depend heavily on direct mappings that achieve only modest accuracy, forcing breeders to conduct large, costly field trials to maintain or marginally improve genetic gain. Models that incorporate intermediate molecular phenotypes such as gene expression can achieve higher predictive fit, but they remain impractical for GS since such data are unavailable at deployment or design time. BINNs overcome this limitation by encoding pathway-level inductive biases and leveraging multi-omics data only during training, while using genotype data alone during inference. Applied to maize gene-expression and multi-environment field-trial data, BINN improves rank-correlation accuracy by up to 56% within and across subpopulations under sparse-data conditions and nonlinearly identifies genes that GWAS/TWAS fail to uncover. With complete domain knowledge for a synthetic metabolomics benchmark, BINN reduces prediction error by 75% relative to conventional neural nets and correctly identifies the most important nonlinear pathway. Importantly, both cases show highly sensitive BINN latent variables correlate with the experimental quantities they represent, despite not being trained on them. This suggests BINNs learn biologically-relevant representations, nonlinear or linear, from genotype to phenotype. Together, BINNs establish a framework that leverages intermediate domain information to improve genomic prediction accuracy and reveal nonlinear biological relationships that can guide genomic selection, candidate gene selection, pathway enrichment, and gene-editing prioritization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。