用大模型知识增强分子属性预测,提升药物研发效率。
Enhancing Molecular Property Prediction with Knowledge from Large Language Models
- 用大模型生成分子知识与向量化代码,融合结构特征
- 在多个数据集上超越现有方法,尤其对冷启动属性有效
- 适合药物发现、分子设计领域研究者参考
分子属性预测是药物研发的关键环节。尽管图神经网络(GNN)和自监督学习已显著推进该任务,但人类先验知识仍不可或缺。本文首次提出将大语言模型(LLM)提取的知识与预训练分子模型的结构特征相结合,以增强分子属性预测(MPP)。我们使用GPT-4o、GPT-4.1和DeepSeek-R1三种先进LLM,通过提示生成与领域相关的知识及可执行的分子向量化代码,构建基于知识的特征,并与结构表示融合。大量实验表明,该方法在多个基准数据集上优于现有方法,证明了知识与结构信息结合的有效性,尤其在低研究度分子属性上表现更优。
原文摘要 · Abstract (English)
Predicting molecular properties is a critical component of drug discovery. Recent advances in deep learning, particularly Graph Neural Networks (GNNs), have enabled end-to-end learning from molecular structures, reducing reliance on manual feature engineering. However, while GNNs and self-supervised learning approaches have advanced molecular property prediction (MPP), the integration of human prior knowledge remains indispensable, as evidenced by recent methods that leverage large language models (LLMs) for knowledge extraction. Despite their strengths, LLMs are constrained by knowledge gaps and hallucinations, particularly for less-studied molecular properties. In this work, we propose a novel framework that, for the first time, integrates knowledge extracted from LLMs with structural features derived from pre-trained molecular models to enhance MPP. Our approach prompts LLMs to generate both domain-relevant knowledge and executable code for molecular vectorization, producing knowledge-based features that are subsequently fused with structural representations. We employ three state-of-the-art LLMs, GPT-4o, GPT-4.1, and DeepSeek-R1, for knowledge extraction. Extensive experiments demonstrate that our integrated method outperforms existing approaches, confirming that the combination of LLM-derived knowledge and structural information provides a robust and effective solution for MPP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。