首个融合分子结构的通用分子大模型,显著提升反应与性质预测性能。
Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization
- 通过结构偏好优化让模型更关注分子图信息
- 在跨任务基准上达顶尖表现,分布外数据上超越旧模型
- 适合药物设计、化学推理等多场景研究者使用
近期大型语言模型(LLMs)在分子任务中取得进展,如化学反应预测和分子性质预测。大规模分子指令微调数据集推动了仅基于序列(如SMILES或SELFIES)的通用分子LLM发展,研究人员正探索结合分子结构信息的多模态方法以进一步提升性能。然而,真正具备多模态能力且覆盖广泛分子任务的通用模型尚未充分研究。我们发现,朴素的下一个词元预测训练会忽略图结构信息,限制模型对分子图的利用。为此,我们提出:(i) 分子结构偏好优化(MolPO),通过优化正确与扰动分子结构之间的偏好关系来促进图信息利用;(ii) 一种改进的图编码器及定制预训练策略,增强MolPO的效果。基于此,我们构建了Mol-LLM,这是首个兼具以下特点的多模态通用模型:(a) 覆盖分子LLM中最广泛的分子任务,(b) 显式利用分子结构信息,(c) 充分利用大规模指令微调。Mol-LLM在最全面的分子LLM基准测试中达到当前最优或可比结果,尤其在反应与性质预测的分布外数据集上显著超越先前通用分子LLM。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have led to models that tackle diverse molecular tasks, such as chemical reaction prediction and molecular property prediction. Large-scale molecular instruction-tuning datasets have enabled sequence-only (e.g., SMILES or SELFIES) generalist molecular LLMs, and researchers are now exploring multimodal approaches that incorporate molecular structural information for further gains. However, a genuinely multimodal, generalist LLM that covers a broad spectrum of molecular tasks has yet to be fully investigated. We observe that naive next token prediction training ignores graph-structural information, limiting an LLM's ability to exploit molecular graphs. To address this, we propose (i) Molecular structure Preference Optimization (MolPO), which facilitates graph usage by optimizing preferences between pairs of correct and perturbed molecular structures, and (ii) an advanced graph encoder with a tailored pre-training strategy to improve the effect of graph utilization by MolPO. Building on these contributions, we introduce Mol-LLM, the first multimodal generalist model that (a) handles a broad spectrum of molecular tasks among molecular LLMs, (b) explicitly leverages molecular-structure information, and (c) takes advantage of extensive instruction tuning. Mol-LLM attains state-of-the-art or comparable results across the most comprehensive molecular-LLM benchmark-even on out-of-distribution datasets for reaction and property prediction, where it surpasses prior generalist molecular LLMs by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。