针对分子模型外推能力评估难题,提出新基准与自适应选源框架。
Rethinking Molecular OOD Generalization via Target-Aware Source Selection

- 基于物理化学特征聚类构建新评测基准,避免语义重叠
- 在多个模型上实现平均误差降低6.2%,最高降11.2%
- 适合关注药物发现中模型泛化能力的研究者
极端分布外(OOD)场景下分子性质的鲁棒预测是人工智能驱动药物研发的关键瓶颈。现有骨架分割评估方法无法消除微观语义重叠,导致模型依赖捷径学习并高估外推能力;传统领域自适应方法在结构剧烈变化时表现差,盲目对齐异构源数据会引入拓扑噪声并引发负向迁移。为此,本文提出基于聚类级划分的分子外推性能评估基准SCOPE-BENCH,以及多源自适应策略优化框架POMA。POMA将知识迁移建模为检索-组合-适配流程:先识别与目标结构接近的标注源骨架作为代理目标;再通过强化学习从指数级候选池中选择最优源子集;最后在宏观拓扑与微观药效团两个尺度上进行域适应。实验表明,先进3D分子模型在SCOPE-BENCH上的预测误差平均上升5.9倍,最高达8.0倍;而POMA在多种骨干网络上实现平均绝对误差降低6.2%,最高降幅11.2%。代码已公开于https://anonymous.4open.science/r/Molecular-OOD-Code-73F6。
原文摘要 · Abstract (English)
Robust prediction of molecular properties under extreme out-of-distribution (OOD) scenarios is a pivotal bottleneck in AI-driven drug discovery. Current scaffold-splitting protocols fail to obstruct microscopic semantic overlap, predisposing models to shortcut learning and overestimating their true extrapolation capability; meanwhile, conventional domain adaptation paradigms suffer under extreme structural shifts, as blindly aligning heterogeneous source libraries injects topological noise and triggers negative transfer. To address these two challenges, scaffold-cluster out-of-distribution performance evaluation benchmark (SCOPE-BENCH), a benchmark built on cluster-level partitioning in an explicit physicochemical descriptor space, is proposed alongside policy optimization for multi-source adaptation (POMA), a framework that formulates knowledge transfer as a retrieve-compose-adapt pipeline: labeled source scaffolds structurally close to the unlabeled target are first identified as proxy targets; a reinforcement learning policy then adaptively selects the optimal source subset from an exponentially large candidate pool; and dual-scale domain adaptation is finally performed at macroscopic topological and microscopic pharmacophore scales. Evaluations show that prediction errors of state-of-the-art 3D molecular models surge by up to 8.0x on SCOPE-BENCH with a mean of 5.9x, while POMA achieves up to an 11.2% reduction in mean absolute error with an average relative improvement of 6.2% across diverse backbone architectures. Code is available at https://anonymous.4open.science/r/Molecular-OOD-Code-73F6.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。