KANEL用可解释网络提升虚拟筛选命中率,更适合早期药物发现。
KANEL: Kolmogorov-Arnold Network Ensemble Learning Enables Early Hit Enrichment in High-Throughput Virtual Screening
- 融合KAN与多种模型,基于不同分子描述符构建集成学习
- 在顶100个预测化合物中,PPV@100显著优于传统模型
- 适合需要高可信度早期候选物的药物研发团队
机器学习模型在化学生物活性预测中被用于从虚拟筛选库中优先选择少量化合物进行实验验证。在此类应用中,通过前N位命中富集率(如PPV@N)评估模型准确性,比传统的全局指标(如AUC)更合适且更具可操作性。本文提出KANEL,一种集成学习流程,结合可解释的柯尔莫哥洛夫-阿诺德网络(KAN)与XGBoost、随机森林及多层感知机模型,这些模型基于互补的分子表征(LillyMol描述符、RDKit衍生描述符和摩根指纹)训练而成。
原文摘要 · Abstract (English)
Machine learning models of chemical bioactivity are increasingly used for prioritizing a small number of compounds in virtual screening libraries for experimental follow-up. In these applications, assessing model accuracy by early hit enrichment such as Positive Predicted Value (PPV) calculated for top N hits (PPV@N) is more appropriate and actionable than traditional global metrics such as AUC. We present KANEL, an ensemble workflow that combines interpretable Kolmogorov-Arnold Networks (KANs) with XGBoost, random forest, and multilayer perceptron models trained on complementary molecular representations (LillyMol descriptors, RDKit-derived descriptors, and Morgan fingerprints).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。