arXiv:2608.30472cs.LG2026-08

提出可复现的分子毒性预测框架,有效避免数据泄漏问题。

ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction

  • 基于图学习与多任务架构,融合化学结构与全局特征
  • 在去泄漏测试集上达0.44的MCC与0.83的AUC
  • 支持不确定性校准与毒性基团发现,适合药物研发

分子毒性预测被广泛用于实验前化合物优先排序,但传统基准性能可能因训练与测试集存在结构相似分子而高估实际效用。本文提出ToxLens,一个可复现的多任务图学习框架,覆盖11个毒性终点,包括Ames致突变性、急性口服毒性、hERG抑制以及Tox21核受体和应激反应检测。该流程整合保守化学清洗、球体排除过滤、泄漏感知UMAP-HDBSCAN划分、并行图编码器与全局特征编码器通过后期拼接融合、温度缩放蒙特卡洛丢弃结合类置信区间预测集、适用域分析及基于SHAP的毒性基团发现(含遮蔽控制)。在去泄漏测试集上,五种子随机投票集成模型取得0.44的马修斯相关系数、0.83的ROC-AUC和0.58的PR-AUC。其表现超越四种基于ECFP4的浅层基线,在全部11个终点上均更优。消融实验表明全局路径重要,后期拼接优于门控与特征线性调制融合方式。类置信区间预测集显示端点间效率差异显著,且判别力与校准度随与训练域相似度提升而增强。在固定发布的Tox21挑战与TDA数据集上重训练,获得具有竞争力但非统一领先的表现。基于SHAP的遮蔽与共识子图挖掘生成模型推导的结构假设,其中44个满足预设反事实标准。

原文摘要 · Abstract (English)

Molecular toxicity prediction is increasingly used to prioritise compounds before experimental testing, but conventional benchmark performance can overstate practical utility when structurally related molecules occur across training and test folds. We introduce ToxLens, a reproducible multi-task graph-learning framework for 11 toxicity endpoints spanning Ames mutagenicity, acute oral toxicity, hERG inhibition, and Tox21 nuclear-receptor and stress-response assays. The workflow combines conservative chemical curation, sphere-exclusion filtering, a leakage-aware UMAP-HDBSCAN split, parallel graph and global-feature encoders joined by late concatenation, temperature-scaled Monte Carlo dropout with conformal-style prediction sets, applicability-domain analysis, and SHAP-guided toxicophore discovery with occlusion controls. On the leakage-controlled test fold, a five-seed soft-voting ensemble achieved a Matthews correlation coefficient score of 0.44, an area under the receiver operating characteristic curve score of 0.83, and an area under the precision-recall curve score of 0.58. It exceeded four ECFP4-based shallow baselines on all 11 endpoints under the same split and validation-based threshold-selection protocol. Controlled ablations showed that the global pathway was important, whereas late concatenation outperformed the tested gated and feature-wise linear modulation fusion variants. Conformal-style prediction sets revealed substantial endpoint-specific variation in set efficiency, and discrimination and calibration improved with similarity to the training domain. Retraining on fixed published Tox21 Challenge and TDA folds produced competitive, but not uniformly state-of-the-art, performance. SHAP-guided occlusion and consensus subgraph mining yielded model-derived structural hypotheses, 44 of which contained at least one occurrence that passed the predefined counterfactual criteria.

毒性预测图神经网络不确定性校准可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。