arXiv:2604.19840cs.LGq-bio.QM2026-04

用图论方法预测分子性质,改进后效果接近甚至超过深度学习。

Graph-Theoretic Models for the Prediction of Molecular Measurements

论文配图:Graph-Theoretic Models for the Prediction of Molecular Measurements
图 1 · 摘自论文原文
  • 基于图指标构建经典模型,逐步加入正则化和特征融合提升性能。
  • 平均R²从0.24提升至0.79,最高提升274%,均显著优于原始模型。
  • 无需GPU、训练快,适合资源有限的研究者使用。

图论方法在分子性质预测中具有简洁、可解释且计算成本低的优势。本文评估了基于外部活性D(G)与内部活性ζ(G)的基准模型在五个MoleculeNet基准数据集上的表现,涵盖生物活性(BACE,1513分子)、脂溶性(LogP合成,14610分子;LogP实验,753分子)、水溶性(ESOL,1128分子)和水合自由能(SAMPL,642分子)。原始模型平均R²仅为0.24,表明泛化能力有限。为此,提出系统性增强框架,依次引入岭回归、额外图描述符、理化性质、梯度提升集成学习、Lasso特征选择及拓扑指数与Morgan指纹的混合方法。增强后模型平均最佳R²达0.79,单项提升幅度165%至274%,所有改进均具统计显著性(p < 0.001)。在相同条件下与图卷积网络对比,增强模型在所有数据集上表现相当或更优。与近期GNN+PGM混合模型相比,本方法在两个数据集上最优,一个数据集并列。整个流程无需GPU,训练时间低于五分钟,仅依赖开源工具,适用于资源受限环境。

原文摘要 · Abstract (English)

Graph-theoretic approaches offer simplicity, interpretability, and low computational cost for molecular property prediction. Among these, the model proposed by Mukwembi and Nyabadza, based on the external activity $D(G)$ and internal activity $ζ(G)$ indices, achieved strong results on a small flavonoid dataset. However, its ability to generalize to larger and chemically diverse datasets has not been tested. This study evaluates the baseline $D(G)$-$ζ(G)$ polynomial model on five benchmark datasets from MoleculeNet, covering biological activity (BACE, 1,513 molecules), lipophilicity (LogP synthetic, 14,610 molecules; LogP experimental, 753 molecules), aqueous solubility (ESOL, 1,128 molecules), and hydration free energy (SAMPL, 642 molecules). The baseline model achieves an average $R^2 = 0.24$, confirming limited transferability. To address this, a systematic enhancement framework is proposed, progressively incorporating Ridge regularization, additional graph descriptors, physicochemical properties, ensemble learning with Gradient Boosting, Lasso feature selection, and a hybrid approach combining topological indices with Morgan fingerprints. The enhanced models raise the average best $R^2$ to 0.79, with individual improvements ranging from 165\% to 274\%. All improvements are statistically significant ($p < 0.001$). A direct comparison with a Graph Convolutional Network under identical experimental conditions shows that the enhanced classical models match or outperform deep learning on all five datasets. Comparison with the recent GNN+PGM hybrid of Djagba et al.\ further confirms competitiveness, with the enhanced models achieving the best results on two datasets and tying on one. The entire framework requires no GPU, trains in under five minutes, and uses only open-source tools, making it accessible for researchers in resource-limited settings.

分子预测图神经网络经典模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。