arXiv:2604.18316q-bio.OTcs.LG2026-04

用机器学习筛选天然药物,加速阿尔茨海默病新药发现。

Predictive Modelling of Natural Medicinal Compounds for Alzheimer disease Using Machine Learning and Cheminformatics

论文配图:Predictive Modelling of Natural Medicinal Compounds for Alzheimer disease Using Machine Learning and Cheminformatics
图 1 · 摘自论文原文
  • 基于化学信息学描述符构建机器学习模型,预测天然化合物抗痴呆活性。
  • 随机森林模型准确率最高,ROC-AUC达0.92以上,优于其他算法。
  • 揭示脂溶性、分子量和极性是影响神经保护作用的关键因素,适合药物研发者参考。

阿尔茨海默病(AD)是一种缺乏特异性治疗手段的神经退行性疾病。天然药物虽具神经保护作用,但高通量发现因实验成本高昂而困难。本研究提出一种机器学习方法,基于来自ChEMBL和PubChem等公共数据库的活性与非活性化合物,利用RDKit计算分子量、脂溶性(LogP)、拓扑极性表面积(TPSA)及氢键描述符等分子描述符。经过数据预处理与特征选择后,构建了随机森林、XGBoost、支持向量机和逻辑回归等多种分类模型,并以准确率、精确率、召回率、F1分数和ROC-AUC进行评估。结果表明,集成学习方法如随机森林表现最佳,其预测准确率与ROC-AUC均超过0.92。特征重要性分析进一步指出,脂溶性、分子量和极性是驱动神经保护活性的关键理化特性。该融合方法展示了将天然产物研究与机器学习结合在痴呆早期药物发现中的潜力,可快速筛选大规模数据集并选出候选分子用于实验验证,显著降低研发成本与时间。

原文摘要 · Abstract (English)

Alzheimer disease (AD) is a neurodegenerative disease that lacks specific treatment options. Natural drugs have displayed neuroprotective effects; however, their high-throughput discovery is challenging because of the expense of experimental testing.The study proposed a machine learning approach to identify the anti-dementia activity of natural compounds based on molecular descriptors obtained from cheminformatics. The study used a set of active and inactive compounds obtained from public databases like ChEMBL and PubChem. Various molecular descriptors, including molecular weight, lipophilicity (LogP), topological polar surface area (TPSA), and hydrogen bonding descriptors, were calculated with RDKit. Data preprocessing and feature selection were applied, followed by the development of several classification models (Random Forest, XGBoost, Support Vector Machines, Logistic Regression) and their evaluation based on accuracy, precision, recall, F1-score and ROC-AUC. The outcome suggests that ensemble techniques, such as Random Forest, delivered the best predictive accuracy and ROC-AUC values. This study also highlights that critical physicochemical descriptors in particular lipophilicity, molecular weight and polarity are important in driving neuroprotective activity as identified by feature importance analysis. The integrated machine learning approach shows the potential of combining natural product research and machine learning in early drug discovery for dementia. They provide a means of rapidly exploring large datasets and selecting candidates for experimental confirmation, thus minimising costs and time in the development of drugs for neurodegenerative diseases.

阿尔茨海默病机器学习天然药物药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。