arXiv:2410.18475cs.AI2024-10被引 1

用跨菌种知识迁移提升代谢物基因关联预测准确率

Gene-Metabolite Association Prediction with Interactive Knowledge Transfer Enhanced Graph for Metabolite Production

  • 构建多菌种代谢图谱,通过预训练模型建立跨图连接
  • 利用跨图链接传播信息,使代谢图谱知识相互增强
  • 在2474种代谢物、1947个基因上实现最高12.3%性能提升

在代谢工程快速发展的背景下,高效精准地识别促进代谢物生产的基因靶点仍面临巨大挑战。传统方法或依赖知识库或基于模型,均因文献规模庞大及基因组尺度代谢模型(GEM)模拟的近似性而耗时费力。为此,我们提出新的任务——基于代谢图的基因-代谢物关联预测,并构建首个基准数据集,包含酿酒酵母(SC)和东方伊萨酵母(IO)的2474种代谢物与1947个基因。该任务因代谢图不完整及不同代谢系统的异质性而困难。为此,我们提出基于代谢图的交互式知识迁移机制(IKT4Meta),通过整合多微生物代谢图知识提升关联预测精度。首先,利用具备外部基因与代谢物知识的预训练语言模型(PLMs)生成跨图连接,缓解异质性影响;其次,以跨图连接为锚点,在各代谢图内部传播链接信息;最后,基于融合多微生物知识的增强代谢图进行基因-代谢物关联预测。在两种生物体上的实验表明,该方法在多种链接预测框架下相较基线最高提升12.3%。

原文摘要 · Abstract (English)

In the rapidly evolving field of metabolic engineering, the quest for efficient and precise gene target identification for metabolite production enhancement presents significant challenges. Traditional approaches, whether knowledge-based or model-based, are notably time-consuming and labor-intensive, due to the vast scale of research literature and the approximation nature of genome-scale metabolic model (GEM) simulations. Therefore, we propose a new task, Gene-Metabolite Association Prediction based on metabolic graphs, to automate the process of candidate gene discovery for a given pair of metabolite and candidate-associated genes, as well as presenting the first benchmark containing 2474 metabolites and 1947 genes of two commonly used microorganisms Saccharomyces cerevisiae (SC) and Issatchenkia orientalis (IO). This task is challenging due to the incompleteness of the metabolic graphs and the heterogeneity among distinct metabolisms. To overcome these limitations, we propose an Interactive Knowledge Transfer mechanism based on Metabolism Graph (IKT4Meta), which improves the association prediction accuracy by integrating the knowledge from different metabolism graphs. First, to build a bridge between two graphs for knowledge transfer, we utilize Pretrained Language Models (PLMs) with external knowledge of genes and metabolites to help generate inter-graph links, significantly alleviating the impact of heterogeneity. Second, we propagate intra-graph links from different metabolic graphs using inter-graph links as anchors. Finally, we conduct the gene-metabolite association prediction based on the enriched metabolism graphs, which integrate the knowledge from multiple microorganisms. Experiments on both types of organisms demonstrate that our proposed methodology outperforms baselines by up to 12.3% across various link prediction frameworks.

代谢工程知识迁移图神经网络基因预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。