线性模型预测脂溶性时误差变大,树模型更准且能破解特征混淆难题。
Diagnosing Heteroskedasticity and Resolving Multicollinearity Paradoxes in Physicochemical Property Prediction
- 用树模型替代线性模型,避免误差随数值增大而放大的问题。
- 树模型在预测脂溶性时R²达0.765,优于传统方法。
- 揭示分子量虽与脂溶性相关弱,却是关键预测因子。
脂溶性(logP)预测在药物发现中至关重要,但现有线性回归模型常违反统计假设,导致性能评估无效。研究分析了来自PubChem、ChEMBL和eMolecules数据库的426,850个生物活性分子,发现预测计算logP值(XLOGP3)的线性模型存在严重异方差性:当logP > 5时,残差方差比中等区域(logP 2–4)高出4.2倍。经典修正方法(加权最小二乘和Box-Cox变换)均未能解决该问题(Breusch-Pagan检验p < 0.0001)。相比之下,树模型(随机森林R²=0.764,XGBoost R²=0.765)对异方差具有天然鲁棒性,且表现更优。SHAP分析揭示了多重共线性悖论:尽管分子量与logP的双变量相关性仅为0.146,但其平均绝对SHAP值高达0.573,是最重要的预测因子,而这一效应被拓扑极性表面积(TPSA)所掩盖。结果表明,标准线性模型在计算脂溶性预测中存在根本缺陷,并为QSAR中集成模型的解释提供了系统框架。
原文摘要 · Abstract (English)
Lipophilicity (logP) prediction remains central to drug discovery, yet linear regression models for this task frequently violate statistical assumptions in ways that invalidate their reported performance metrics. We analyzed 426,850 bioactive molecules from a rigorously curated intersection of PubChem, ChEMBL, and eMolecules databases, revealing severe heteroskedasticity in linear models predicting computed logP values (XLOGP3): residual variance increases 4.2-fold in lipophilic regions (logP greater than 5) compared to balanced regions (logP 2 to 4). Classical remediation strategies (Weighted Least Squares and Box-Cox transformation) failed to resolve this violation (Breusch-Pagan p-value less than 0.0001 for all variants). Tree-based ensemble methods (Random Forest R-squared of 0.764, XGBoost R-squared of 0.765) proved inherently robust to heteroskedasticity while delivering superior predictive performance. SHAP analysis resolved a critical multicollinearity paradox: despite a weak bivariate correlation of 0.146, molecular weight emerged as the single most important predictor (mean absolute SHAP value of 0.573), with its effect suppressed in simple correlations by confounding with topological polar surface area (TPSA). These findings demonstrate that standard linear models face fundamental challenges for computed lipophilicity prediction and provide a principled framework for interpreting ensemble models in QSAR applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。