用知识图谱分解蛋白互作路径,提升药物靶点发现的可解释性。
GraPPI: A Retrieve-Divide-Solve GraphRAG Framework for Large-scale Protein-protein Interaction Exploration
- 基于知识图谱将蛋白互作路径拆解为子任务分析
- 支持大规模互作路径探索,生成多条潜在治疗靶点
- 强调结果可解释性,适合药物研发人员使用
药物研发对公共卫生至关重要。研究认为抑制蛋白质错误折叠可延缓疾病进程,因此靶点识别(Target ID)成为关键。尽管大语言模型(LLMs)和检索增强生成(RAG)框架加速了药物研发,但如何将其整合为连贯工作流仍具挑战。我们通过用户调研发现:1)模型应基于初始蛋白提供多个蛋白-蛋白互作(PPI)及具治疗潜力的候选蛋白;2)需给出互作关系及其解释以增强理解。现有方法存在三大局限:语义模糊、缺乏可解释性、检索单元过短。为此,我们提出GraPPI——一种基于大规模知识图谱(KG)的“检索-分割-求解”代理管道RAG框架,通过将完整PPI通路分析分解为聚焦于互作边的子任务,支持大规模PPI信号通路探索,以揭示治疗影响。
原文摘要 · Abstract (English)
Drug discovery (DD) has tremendously contributed to maintaining and improving public health. Hypothesizing that inhibiting protein misfolding can slow disease progression, researchers focus on target identification (Target ID) to find protein structures for drug binding. While Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) frameworks have accelerated drug discovery, integrating models into cohesive workflows remains challenging. We conducted a user study with drug discovery researchers to identify the applicability of LLMs and RAGs in Target ID. We identified two main findings: 1) an LLM should provide multiple Protein-Protein Interactions (PPIs) based on an initial protein and protein candidates that have a therapeutic impact; 2) the model must provide the PPI and relevant explanations for better understanding. Based on these observations, we identified three limitations in previous approaches for Target ID: 1) semantic ambiguity, 2) lack of explainability, and 3) short retrieval units. To address these issues, we propose GraPPI, a large-scale knowledge graph (KG)-based retrieve-divide-solve agent pipeline RAG framework to support large-scale PPI signaling pathway exploration in understanding therapeutic impacts by decomposing the analysis of entire PPI pathways into sub-tasks focused on the analysis of PPI edges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。