arXiv:2412.15790q-bio.QMcs.AI2024-12被引 3

用大模型生成的序列嵌入增强图神经网络,提升多组学数据分析精度。

GraphSeqLM: A Unified Graph Language Framework for Omic Graph Learning

  • 融合大模型生成的生物序列嵌入与图结构信息
  • 在多组学数据上预测准确率超越现有方法
  • 适合精准医学中复杂疾病分析的研究者

多组学数据整合对理解复杂疾病至关重要,但其高维性和噪声带来了巨大挑战。图神经网络(GNN)虽能有效分析大规模信号通路和蛋白质互作网络,但在捕捉复杂生物关系时表达能力有限。为此,我们提出图序列语言模型(GraphSeqLM),通过大型语言模型(LLMs)生成的DNA、RNA和蛋白质序列嵌入,增强GNN的特征表示能力。这些嵌入编码了序列的结构与生物学特性,使GNN能够更全面地分析样本特异性多组学数据。通过融合拓扑结构、序列衍生信息与生物功能,GraphSeqLM在预测任务中表现更优,显著超越现有方法,为精准医学中的多组学数据整合提供了新路径。

原文摘要 · Abstract (English)

The integration of multi-omic data is pivotal for understanding complex diseases, but its high dimensionality and noise present significant challenges. Graph Neural Networks (GNNs) offer a robust framework for analyzing large-scale signaling pathways and protein-protein interaction networks, yet they face limitations in expressivity when capturing intricate biological relationships. To address this, we propose Graph Sequence Language Model (GraphSeqLM), a framework that enhances GNNs with biological sequence embeddings generated by Large Language Models (LLMs). These embeddings encode structural and biological properties of DNA, RNA, and proteins, augmenting GNNs with enriched features for analyzing sample-specific multi-omic data. By integrating topological, sequence-derived, and biological information, GraphSeqLM demonstrates superior predictive accuracy and outperforms existing methods, paving the way for more effective multi-omic data integration in precision medicine.

多组学图神经网络大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。