arXiv:2603.01537cs.AIq-bio.BM2026-03

不用化学结构也能准确预测药物重定位,靠的是靶点和网络拓扑信息。

Pharmacology Knowledge Graphs: Do We Need Chemical Structure for Drug Repurposing?

  • 用药物靶点和网络拓扑代替化学结构,提升预测效率
  • 去掉化学结构编码后,性能微升(PR-AUC从0.5631到0.5785)
  • 适合关注模型简化与可解释性的药物研发人员

模型复杂度、数据量与特征模态对知识图谱驱动药物重定位的贡献在严格时间验证下仍不明确。我们基于ChEMBL 36构建了包含5,348个实体(3,127种药物、1,156个蛋白、1,065个适应症)的药理知识图谱,并采用严格的时序划分:训练数据截至2022年,测试数据覆盖2023–2025年,同时引入来自失败试验和临床试验的生物学验证硬负例。在五个知识图嵌入模型和一个含344万参数的图注意力编码器与ESM-2蛋白嵌入的GNN上进行基准测试。通过规模从0.78到9.75百万参数、数据量从25%到100%的扩展实验,以及特征消融分析,发现移除基于图注意力的药物结构编码器,仅保留拓扑嵌入与ESM-2蛋白特征,使药物-蛋白PR-AUC从0.5631升至0.5785,显存占用从5.30 GB降至353 MB。用摩根指纹替代药物编码进一步降低性能,表明显式化学结构表示可能有害。模型参数超过244万后收益递减,而增加训练数据持续提升性能。外部验证确认前14个新预测中有6个为已知治疗指征。结果表明,仅凭靶点中心信息与药物网络拓扑即可高精度预测药理行为,无需显式化学结构表示。

原文摘要 · Abstract (English)

The contributions of model complexity, data volume, and feature modalities to knowledge graph-based drug repurposing remain poorly quantified under rigorous temporal validation. We constructed a pharmacology knowledge graph from ChEMBL 36 comprising 5,348 entities including 3,127 drugs, 1,156 proteins, and 1,065 indications. A strict temporal split was enforced with training data up to 2022 and testing data from 2023 to 2025, together with biologically verified hard negatives mined from failed assays and clinical trials. We benchmarked five knowledge graph embedding models and a standard graph neural network with 3.44 million parameters that incorporates drug chemical structure using a graph attention encoder and ESM-2 protein embeddings. Scaling experiments ranging from 0.78 to 9.75 million parameters and from 25 to 100 percent of the data, together with feature ablation studies, were used to isolate the contributions of model capacity, graph density, and node feature modalities. Removing the graph attention based drug structure encoder and retaining only topological embeddings combined with ESM-2 protein features improved drug protein PR-AUC from 0.5631 to 0.5785 while reducing VRAM usage from 5.30 GB to 353 MB. Replacing the drug encoder with Morgan fingerprints further degraded performance, indicating that explicit chemical structure representations can be detrimental for predicting pharmacological network interactions. Increasing model size beyond 2.44 million parameters yielded diminishing returns, whereas increasing training data consistently improved performance. External validation confirmed 6 of the top 14 novel predictions as established therapeutic indications. These results show that drug pharmacological behavior can be accurately predicted using target-centric information and drug network topology alone, without requiring explicit chemical structure representations.

药物重定位知识图谱拓扑建模无结构预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。