让未知分子在生物医学图谱中找到位置,提升药物发现效率。
MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring

- 通过多尺度结构锚定,将未注册分子关联到知识图谱。
- 多跳推理准确率提升至0.876,外部分子召回率达0.269。
- 无需训练即可推理,结果可追溯结构依据和证据来源。
生物医学知识图谱(KG)加速药物发现,但传统流程假设查询分子已存在于图中,导致未注册分子无法连接。本文提出MolBioKG,解决这一冷启动问题——即‘图外分子’难题。该系统为双层架构,通过多分辨率结构锚定,将274万分子(以骨架、片段、官能团、指纹表示)与960万条边的KG关联。仅需SMILES字符串,MolBioKG即可检索结构相似的图实体,并在其生物医学邻域中遍历,无需任务特定训练。系统包含两种推理机制:基于倒数排名融合的静态多锚点检索,以及基于工具使用的自适应大模型策略(Adapt-KG)。在图内链接恢复、复杂多跳推理及图外泛化测试中表现优异,显著提升性能:多跳推理的Hits@10从0.585升至0.876,图外目标召回率从0.145增至0.269。所有预测均保留可追溯的结构锚点和源自知识图谱的证据。
原文摘要 · Abstract (English)
Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities, leaving unregistered molecules disconnected. We address this cold-start challenge, termed the out-of-graph molecule problem, by introducing MolBioKG. This two-layer system grounds unseen molecules in biomedical evidence via multi-resolution structural anchoring. It connects an index of 2.74 million molecules (represented by scaffolds, fragments, functional groups, and fingerprints) to a 9.6-million-edge KG. Given only a SMILES string, MolBioKG retrieves structurally related graph entities and traverses their biomedical neighborhoods without task-specific training. It features two inference mechanisms: static multi-anchor retrieval using Reciprocal Rank Fusion, and Adapt-KG, a tool-using LLM policy for adaptive traversal. Evaluated across in-graph link recovery, complex multi-hop reasoning, and out-of-graph generalization, MolBioKG outperforms strong baselines. Notably, it raises Hits@10 from 0.585 to 0.876 in multi-hop reasoning and out-of-graph target recall from 0.145 to 0.269, all while ensuring predictions retain traceable structural anchors and source-attributed KG evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。