用多上下文采样提升药物靶点预测,不依赖负样本且更准确。
MuCoS: Efficient Drug Target Discovery via Multi Context Aware Sampling in Knowledge Graphs
- 通过高密度邻居采样捕捉图结构模式,融合BERT上下文信息。
- 在KEGG50k数据集上MRR提升13%,药物靶点预测准确率提高6%。
- 适合需高效预测新药靶点的生物医药研究者使用。
准确预测药物靶点相互作用对加速药物研发和揭示复杂生物机制至关重要。本文将药物靶点预测建模为异构生物医学知识图谱(KG)上的链接预测任务,该图谱整合了药物、蛋白质、疾病、通路等实体。传统KG嵌入方法如TransE和ComplEx SE受限于计算密集型负样本采样,且难以泛化到未见的药物-靶点对。为此,我们提出多上下文感知采样(MuCoS)框架,优先选择高密度邻居以捕捉显著结构模式,并结合BERT生成的上下文嵌入。通过统一结构与文本模态并有选择地采样高信息量模式,MuCoS无需负样本采样,显著降低计算开销,同时提升对新型药物-靶点关联及药物靶点的预测精度。在KEGG50k数据集上的大量实验表明,MuCoS优于现有最优基线,在预测任意关系时平均倒数排名(MRR)最高提升13%,在专门的药物靶点关系预测中提升6%。
原文摘要 · Abstract (English)
Accurate prediction of drug target interactions is critical for accelerating drug discovery and elucidating complex biological mechanisms. In this work, we frame drug target prediction as a link prediction task on heterogeneous biomedical knowledge graphs (KG) that integrate drugs, proteins, diseases, pathways, and other relevant entities. Conventional KG embedding methods such as TransE and ComplEx SE are hindered by their reliance on computationally intensive negative sampling and their limited generalization to unseen drug target pairs. To address these challenges, we propose Multi Context Aware Sampling (MuCoS), a novel framework that prioritizes high-density neighbours to capture salient structural patterns and integrates these with contextual embeddings derived from BERT. By unifying structural and textual modalities and selectively sampling highly informative patterns, MuCoS circumvents the need for negative sampling, significantly reducing computational overhead while enhancing predictive accuracy for novel drug target associations and drug targets. Extensive experiments on the KEGG50k dataset demonstrate that MuCoS outperforms state-of-the-art baselines, achieving up to a 13\% improvement in mean reciprocal rank (MRR) in predicting any relation in the dataset and a 6\% improvement in dedicated drug target relation prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。