用AI预测未发现的科学关系,助力药物研发新突破
Hakken: Predicting future discoveries to fill the gaps in today's knowledge

- 基于时序知识图谱与大模型语义,预测科学概念间新关系
- 在衰老研究中生成150万条高置信度假设,3条经实验证实
- 可辅助生物学家发现潜在药物靶点,适合医学与交叉领域研究
我们提出Hakken,一个不依赖特定领域的知识预测与解释系统,通过构建基于海量科研文献提取的时序知识图谱,并融合大语言模型的语义知识,预测尚未记录的科学概念间新型关系。该系统采用基于Transformer的预测模型,识别并定义这些新关系的存在与类型,再调用模型无关的解释框架提供支撑信息,帮助科学家评估预测合理性。虽为通用系统,我们以生物医学领域为例展示了其实际能力:其预测模型在时间感知的多标签关系预测任务上达到新基准,且在历史数据中保持长期一致性。系统共生成150万条高于置信阈值的衰老相关假设,经生物学家定性验证后,推进三条进入湿实验验证。其中两条具有重要潜力:首次揭示了TP53与BAMBI、RAF1与TNF间的未被记录相互作用,为药物发现与老药新用提供新线索。
原文摘要 · Abstract (English)
We present Hakken, a domain-agnostic prediction and explanation system performing knowledge prediction, i.e., growing scientific knowledge by establishing novel relationships, ones that are not limited to the deductive hull of previous knowledge. Hakken uses a transformer-based prediction model built on temporal sequences of knowledge graphs extracted from vast bodies of research publications, fused with an LLM's semantic knowledge, to predict the presence and define the type of as-yet undocumented relationships between scientific concepts. It then calls a model-agnostic explanation framework to provide accompanying information for each prediction that allows scientists to evaluate the suggested new relationship. While general purpose, we demonstrate Hakken's practical capabilities by applying it to the biomedical domain. There, Hakken's prediction model establishes a new benchmark for time-aware multi-label relation prediction, and we show that the model's output stays coherent and informative over extended time spans in historic data. In addition, we scored 1.5 million above-confidence-threshold hypotheses related to aging, qualitatively validated batches of these predictions with biologists and progressed three of them for empirical validation in wet-lab. Two predictions with potentially significant impact in the context of drug discovery and repurposing were confirmed, introducing previously undocumented interactions between TP53 and BAMBI, and between RAF1 and TNF, to biomedical science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。