用多指纹与文本-分子对齐模型,高效优化分子结构
MolLIBRA: Genetic Molecular Optimization with Multi-Fingerprint Surrogates and Text-Molecule Aligned Critic
- 结合多种分子指纹和文本-分子编码器预排候选分子
- 在1000次评估内,14个任务达最优Top-10 AUC
- 适合药物设计等需少样本优化的分子研发场景
我们研究在有限查询预算下的高效分子优化问题。提出MolLIBRA(多模态与语言集成的贝叶斯与进化优化框架),一种基于遗传算法的框架,在调用真实评估器前利用多个评判器对候选分子进行预排序:(i) 基于多种分子指纹的高斯过程(GP)代理集成;(ii) 预训练的文本-分子对齐编码器CLAMP。GP集成可自适应选择任务相关指纹,而CLAMP通过比较分子与文本嵌入相似性,从任务描述中提供零样本评分信号。在1000次评估预算(PMO-1K)的实用分子优化基准上,采用语言模型生成器的MolLIBRA-L版本在22项任务中的14项取得最佳Top-10 AUC,且跨任务总Top-10 AUC最高。
原文摘要 · Abstract (English)
We study sample-efficient molecular optimization under a limited budget of oracle evaluations. We propose MolLIBRA (MultimOdaLity and Language Integrated Bayesian and evolutionaRy optimizAtion), a genetic algorithm based framework that pre-ranks candidate molecules using multiple critics before oracle calls: (i) an ensemble of Gaussian process (GP) surrogates defined over multiple molecular fingerprints and (ii) a pretrained text-molecule aligned encoder CLAMP. The GP ensemble enables adaptive selection of task-appropriate fingerprints, while CLAMP provides a zero-shot scoring signal from task descriptions by measuring the similarity between molecular and text embeddings. On the Practical Molecular Optimization (PMO) benchmark with a budget of 1,000 evaluations (PMO-1K), MolLIBRA-L, our variant with a language-model-based candidate generator, attains the best Top-10 AUC on 14/22 tasks and the highest overall sum of Top-10 AUC across tasks among prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。