新评估方法发现现有药物模型在陌生结构上表现严重下滑。
Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models

- 用结构前沿分拆法替代传统骨架分隔,更严格测试模型泛化能力。
- 在6个药物性质任务中,错误率中位数飙升87.0%,最高达246.0%。
- 不同分子视角和终点影响预测可靠性,提示评估需谨慎设计。
分子性质模型常通过保留Bemis-Murcko骨架进行评估,但骨架仅是化学陌生性的一种定义。本文提出一种无标签的结构前沿分拆方法,保留最稀疏且理化性质最远的骨架组,并在六个公开的实验或人工标注的ADMET任务上进行评估。与70/10/20骨架控制组(相同开链分组)相比,前沿分拆使等权重主误差中位数上升87.0%,偏斜敏感均值达130.3%(描述性任务/种子自举区间:52.1–246.0%)。去除血脑屏障(BBB)后均值降至75.9%,该终点在前沿下预测排名反转。消息传递图网络控制组仍存在显著差距(四任务均值82.8%),且未出现反转,说明低容量头部非根本原因。测试多视图前沿风险外推(MV-FREX)——四种分子视图的计数调整尾部风险惩罚,作为可验证探针,其相对于经验风险最小化的误差变化仅0.16%(区间-0.43–0.84%),图网络为-1.9%;三种固定鲁棒惩罚对照亦无显著效果。相较已发表的Lo-Hi与DataSAIL分拆,前沿分拆平均使误差更高,但无一始终最难。对31,561个海洋天然产物的审计显示,外部样本状态与历史预测一致性取决于分子视图、终点及教师覆盖范围。分拆构建与标签来源本身即为重要评估约束,所测试的训练惩罚无法缓解观察到的前沿失败。
原文摘要 · Abstract (English)
Molecular property models are commonly evaluated by holding out Bemis-Murcko scaffolds, yet a scaffold identifier is only one notion of chemical unfamiliarity. We introduce a label-free structural-frontier split that reserves the sparsest and most physicochemically remote scaffold groups, and evaluate it on six public experimental or curated ADMET tasks. Against a 70/10/20 scaffold control with identical acyclic grouping, the frontier inflates equally weighted primary error with a taskwise median of 87.0% and a skew-sensitive mean of 130.3% (descriptive task/seed bootstrap interval, 52.1-246.0%). The mean falls to 75.9% once BBB is removed; that endpoint is the one whose score ranking inverts at the frontier. A message-passing graph-network control still shows a large gap (mean 82.8% over four tasks) and does not invert, so a low-capacity head does not explain the effect. We also test Multi-View Frontier Risk Extrapolation (MV-FREX), a count-adjusted tail-risk penalty over four molecular views, and treat it as a falsifiable probe. It changes normalized frontier error by only 0.16% relative to empirical risk minimization for the perceptron head (interval, -0.43-0.84%) and by -1.9% for the graph network; three fixed robust-penalty controls are likewise inconclusive. Against the published Lo-Hi and DataSAIL splitters, the frontier inflates error more on average, though no split is uniformly hardest. An audit of 31,561 marine natural products further shows that OOD status and agreement with legacy ADMET predictions depend on the molecular view, endpoint, and teacher coverage. Split construction and label provenance are important evaluation constraints in their own right, and the tested training penalties do not resolve the frontier failures we observe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。