用分子隐空间优化任务增强表格大模型,提升生成式设计效率
In-Context Learning for Latent Space Bayesian Optimization

- 在分子VAE隐空间构造合成优化任务,补充预训练
- 新模型在分子优化基准上表现优异,超越基线
- 适合需要高效生成设计的药物分子研发人员
贝叶斯优化(BO)是高效设计的核心工具,而隐空间贝叶斯优化(LSBO)将其扩展至分子、蛋白质等结构化对象。与此同时,如TabPFN和TabICL等表格基础模型通过大规模合成预训练,在回归任务中达到顶尖性能,正被广泛用作BO代理模型。然而,其贝叶斯行为依赖于预训练分布的构成。在LSBO中,从隐码到目标值的映射与现有上下文学习模型所训练的回归任务存在显著差异。为此,本文在表格基础模型的预训练阶段加入基于分子VAE隐空间的合成优化任务,继续预训练目标包含正则项,保持原始检查点的广泛回归先验,避免对适应任务过度特化。在分子优化的留出测试集上,该模型展现出强性能,验证了针对LSBO特化适配的必要性。
原文摘要 · Abstract (English)
Bayesian optimization (BO) is a central tool for sample-efficient design, and latent-space Bayesian optimization (LSBO) extends it to structured objects such as molecules and proteins. In parallel, tabular foundation models such as TabPFN and TabICL now achieve state-of-the-art regression performance and are increasingly used as BO surrogates. Because their Bayesian behavior is induced by large synthetic pretraining collections, the composition of this pretraining distribution is crucial. LSBO creates a distinctive mismatch: the induced map from latent code to objective value differs markedly from the regression tasks used to train current in-context models. We address this mismatch by complementing the pretraining stage of tabular foundation model surrogates with synthetic optimization tasks defined on the latent space of a molecular VAE. The continued-pretraining objective features a regularizer that anchors the model to the original checkpoint, preserving its broad regression prior while avoiding overspecialization to the adaptation tasks. On held-out molecular optimization benchmarks, the resulting model achieves strong performance, supporting the relevance of LSBO-specific adaptation for in-context surrogates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。