让大模型学会多属性分子优化,无需重新训练就能应对新任务。
GeLLMO: Generalizing Large Language Models for Multi-property Molecule Optimization
- 用指令微调数据集训练大模型,实现多属性分子优化。
- 在10个任务中表现超越现有方法,零样本泛化能力更强。
- 适合药物研发、材料设计等需要快速响应新目标的场景。
尽管近期取得进展,大多数分子优化的计算方法仍局限于单属性或双属性优化,且可扩展性和对新任务的泛化能力较差。与此同时,大语言模型(LLMs)展现出出色的跨领域任务泛化能力。为验证LLMs在分子优化中的潜力,我们提出首个高质量的指令微调数据集MuMOInstruct,专门针对复杂的多属性分子优化任务。基于该数据集,我们开发了GeLLMO系列指令微调的LLM,用于分子优化。在5个域内和5个域外任务上的广泛评估表明,GeLLMO consistently优于当前最优基线。此外,GeLLMO在未见任务上表现出卓越的零样本泛化能力,显著优于强大的闭源大模型。这种强泛化能力展示了GeLLMO作为分子优化基础模型的巨大潜力,可在无需资源密集型重训练的情况下应对新型优化任务。MuMOInstruct数据集、模型及代码已开源:https://github.com/ninglab/GeLLMO。
原文摘要 · Abstract (English)
Despite recent advancements, most computational methods for molecule optimization are constrained to single- or double-property optimization tasks and suffer from poor scalability and generalizability to novel optimization tasks. Meanwhile, Large Language Models (LLMs) demonstrate remarkable out-of-domain generalizability to novel tasks. To demonstrate LLMs' potential for molecule optimization, we introduce MuMOInstruct, the first high-quality instruction-tuning dataset specifically focused on complex multi-property molecule optimization tasks. Leveraging MuMOInstruct, we develop GeLLMOs, a series of instruction-tuned LLMs for molecule optimization. Extensive evaluations across 5 in-domain and 5 out-of-domain tasks demonstrate that GeLLMOs consistently outperform state-of-the-art baselines. GeLLMOs also exhibit outstanding zero-shot generalization to unseen tasks, significantly outperforming powerful closed-source LLMs. Such strong generalizability demonstrates the tremendous potential of GeLLMOs as foundational models for molecule optimization, thereby tackling novel optimization tasks without resource-intensive retraining. MuMOInstruct, models, and code are accessible through https://github.com/ninglab/GeLLMO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。