揭示蛋白降解药物预测中性能差距主因,发现实验室差异是关键瓶颈
Decomposing the Generalization Gap in PROTAC Activity Prediction: Variance Attribution and the Inter-Laboratory Ceiling

- 通过分解泛化差距,识别出跨实验室测量差异是性能下降主因
- 不同实验条件导致的误差使模型在新靶点上表现下降至AUROC 0.67
- 提出校准与少样本重训练方案,可将性能提升至0.705,适合药物研发人员参考
机器学习预测生物活性常出现随机划分与留一靶标外(LOTO)之间的显著泛化差距。本文以靶向蛋白降解剂(PROTAC)为案例,研究发现:随机划分下预测性能(AUROC 0.85–0.91)远高于LOTO协议下的表现(约0.67)。前者依赖同靶点内插,后者衡量真实新靶点预测能力。我们分解该差距,发现跨实验室测量变异是主导因素,其贡献达0.124 AUROC,远超二值化阈值选择的0.05贡献。在八种架构及高达30亿参数的ESM-2模型中,LOTO AUROC均稳定在0.67左右,即使去重或超参优化也无法突破此天花板。单种子配置在多种子验证中损失0.161 AUROC,符合选择偏差理论预测。结合少样本(k=5)分靶重训与ADMET特征,65靶点LOTO AUROC从0.668升至0.705,后处理Platt校准恢复至0.05内良好校准。论文发布PROTAC-Bench数据集(10,748条测量、173个靶点、65个LOTO折),以及框架、校准协议与代码。
原文摘要 · Abstract (English)
Machine-learning predictors of biochemical activity often exhibit large random-split-to-leave-one-target-out generalisation gaps that have been documented but not decomposed. We frame this as an evaluation-science question and use targeted protein degradation as the empirical test bed. PROTACs (proteolysis-targeting chimeras) are heterobifunctional small molecules that induce targeted protein degradation, with more than forty candidates currently in clinical trials; published predictors report AUROC of 0.85 to 0.91 under random-split cross-validation, while the leave-one-target-out (LOTO) protocol of Ribes et al. reduces performance to approximately 0.67. Random splits reward within-target interpolation, whereas LOTO measures the novel-target prediction that de-novo design depends on. We decompose this gap and identify inter-laboratory measurement variance as the dominant component, anchored by a within-target cross-laboratory cascade bounding the inter-laboratory contribution at 0.124 AUROC, well above the 0.05 contribution from binarisation-threshold choice. Across eight published architectures and ESM-2 protein language models up to 3B parameters, LOTO AUROC plateaus near 0.67, with a comparable plateau under SMILES-level deduplication; a 21-dimensional 2000-trial hyperparameter optimisation cannot break this ceiling, and the rank-1 single-seed configuration regresses by 0.161 AUROC under multi-seed validation, matching a closed-form selection-bias prediction (Bailey and Lopez de Prado, 2014). Few-shot k=5 stratified per-target retraining combined with ADMET features lifts 65-target LOTO AUROC from 0.668 to 0.7050, and post-hoc Platt scaling recovers raw output to within the 0.05 well-calibrated threshold. We release PROTAC-Bench (10,748 measurements, 173 targets, 65 LOTO folds), the variance-decomposition framework, the per-target calibration protocol, and the evaluation code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。