arXiv:2505.23987cs.LGcs.AI2025-05EMNLP被引 5

让大模型精准优化药物分子多属性,且能零样本适应新任务。

Large Language Models for Controllable Multi-property Multi-objective Molecule Optimization

  • 构建首个聚焦属性特异性目标的指令微调数据集
  • 在10项任务中成功率最高提升126%,显著优于基线
  • 零样本泛化能力强,适合真实药物设计场景

在现实药物设计中,分子优化需选择性提升多个分子属性至药理相关水平,同时保持已达标属性不变。然而,现有计算方法和指令微调大模型难以捕捉这种精细的属性特定目标,限制了实际应用。为此,我们提出 C-MuMOInstruct,首个专注于多属性优化且具备明确属性特异性目标的指令微调数据集。基于该数据集,我们开发 GeLLMO-Cs 系列指令微调大模型,可实现靶向属性优化。在5个分布内和5个分布外任务上的实验表明,GeLLMO-Cs 持续优于强基线,成功率最高提升126%。尤为突出的是,该模型展现出优异的零样本泛化能力,能应对新优化任务和未见指令。这为构建支持真实、多样化属性特异性优化的基础大模型迈出了关键一步。C-MuMOInstruct 数据集与代码已开源:https://github.com/ninglab/GeLLMO-C。

原文摘要 · Abstract (English)

In real-world drug design, molecule optimization requires selectively improving multiple molecular properties up to pharmaceutically relevant levels, while maintaining others that already meet such criteria. However, existing computational approaches and instruction-tuned LLMs fail to capture such nuanced property-specific objectives, limiting their practical applicability. To address this, we introduce C-MuMOInstruct, the first instruction-tuning dataset focused on multi-property optimization with explicit, property-specific objectives. Leveraging C-MuMOInstruct, we develop GeLLMO-Cs, a series of instruction-tuned LLMs that can perform targeted property-specific optimization. Our experiments across 5 in-distribution and 5 out-of-distribution tasks show that GeLLMO-Cs consistently outperform strong baselines, achieving up to 126% higher success rate. Notably, GeLLMO-Cs exhibit impressive 0-shot generalization to novel optimization tasks and unseen instructions. This offers a step toward a foundational LLM to support realistic, diverse optimizations with property-specific objectives. C-MuMOInstruct and code are accessible through https://github.com/ninglab/GeLLMO-C.

分子优化大模型药物设计零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。