arXiv:2509.11493cs.LGq-bio.QM2025-09ICML

用深度聚类与图神经网络挖掘药物新用途,提升发现效率。

Drug Repurposing Using Deep Embedded Clustering and Graph Neural Networks

  • 先用无监督聚类压缩多组学数据,再用图网络预测药病关联。
  • 模型准确率达90.1%,曲线下面积0.960,477个高置信度关联被识别。
  • 适合药理研究者和机器学习在医疗应用中的探索者。

药物重定位传统上因成本过高难以推进,而现代机器学习虽能捕捉复杂生化关系,但多数研究依赖已知药病相似性的简化数据集。本文提出一种融合无监督深度嵌入聚类与有监督图神经网络链接预测的机器学习流程,从多组学数据中发现新的药物-疾病关联。通过自编码器与聚类训练,将9022种独特药物降维至35个簇,平均轮廓系数达0.8550。图神经网络表现优异,预测准确率为0.901,受试者工作特征曲线下面积为0.960,F1分数为0.901。生成了包含477个每簇链接概率超99%的排序列表。该研究可为跨疾病领域提供新药病关联线索,并推动机器学习在药物重定位中的理解与应用。

原文摘要 · Abstract (English)

Drug repurposing has historically been an economically infeasible process for identifying novel uses for abandoned drugs. Modern machine learning has enabled the identification of complex biochemical intricacies in candidate drugs; however, many studies rely on simplified datasets with known drug-disease similarities. We propose a machine learning pipeline that uses unsupervised deep embedded clustering, combined with supervised graph neural network link prediction to identify new drug-disease links from multi-omic data. Unsupervised autoencoder and cluster training reduced the dimensionality of omic data into a compressed latent embedding. A total of 9,022 unique drugs were partitioned into 35 clusters with a mean silhouette score of 0.8550. Graph neural networks achieved strong statistical performance, with a prediction accuracy of 0.901, receiver operating characteristic area under the curve of 0.960, and F1-Score of 0.901. A ranked list comprised of 477 per-cluster link probabilities exceeding 99 percent was generated. This study could provide new drug-disease link prospects across unrelated disease domains, while advancing the understanding of machine learning in drug repurposing studies.

药物重定位图神经网络多组学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。