arXiv:2607.06224cs.LG2026-07中稿 · ICML

用知识图谱和多模态模型预测微生物发酵产量,提升代谢工程效率。

Canopy: A Heterograph Foundation Model for Metabolic Engineering

  • 构建包含690万节点的异构生物知识图谱,融合基因、蛋白、代谢物等13类实体
  • 预训练模型在发酵产量预测上达到R²=0.41,显著优于传统方法的0.24
  • 适合代谢工程、合成生物学研究者,尤其关注高通量数据建模的应用

设计能以商业可行产量生产高价值化学品的微生物菌株,仍是代谢工程的核心挑战。现有计算方法或依赖无法学习实验数据的化学计量约束模型,或使用手工特征的表格机器学习,忽视了生物知识的关联结构。我们提出Canopy,一个异构图基础模型,将十个公开及私有数据源整合为统一知识图谱(KG),包含690万节点,覆盖13种类型与34种边类型,涵盖基因、蛋白、代谢物、反应、通路、菌株及发酵实验。节点特征通过领域专用基础模型(ESM-2用于蛋白序列,MoLFormer用于化学SMILES,PubMedBERT用于生物医学文本)编码,实现单图内多模态表示。我们使用四种自监督目标(链接预测、掩码节点建模、距离预测、对比实验聚类)对异构图变换器(HGT)进行预训练,结合SignNet位置编码、跳跃知识聚合与虚拟节点,并通过学习到的同方差不确定性加权平衡。在下游发酵产量预测任务中,冻结的Canopy嵌入配合轻量探测器达到R²=0.41,优于表格基线(最佳R²=0.24)及同质GNN变体。

原文摘要 · Abstract (English)

Designing microbial strains that produce high-value chemicals at commercially viable titers remains a central challenge in metabolic engineering. Existing computational approaches either rely on stoichiometric constraint-based models that cannot learn from experimental data, or apply tabular machine learning to hand-crafted features that discard the relational structure of biological knowledge. We present Canopy, a heterogeneous graph foundation model that integrates ten public and proprietary data sources into a unified knowledge graph (KG) of 6.9M nodes across 13 types and 34 edge types, covering genes, proteins, metabolites, reactions, pathways, strains, and fermentation experiments. Node features are encoded through domain-specific foundation models (ESM-2 for protein sequences, MoLFormer for chemical SMILES, and PubMedBERT for biomedical text), yielding a multi-modal representation within a single graph. We pretrain a Heterogeneous Graph Transformer (HGT) augmented with SignNet positional encodings, Jumping Knowledge aggregation, and virtual nodes using four self-supervised objectives (link prediction, masked node modelling, distance prediction, and contrastive experiment clustering), balanced via learned homoscedastic uncertainty weighting. On the downstream task of fermentation titer prediction, frozen Canopy embeddings achieve $R^{2} = 0.41$ with a lightweight probe, outperforming tabular baselines (best $R^{2} = 0.24$) and homogeneous GNN variants.

代谢工程知识图谱多模态图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。