用双任务学习提升药物合成预测准确率
Enhancing Chemical Reaction and Retrosynthesis Prediction with Large Language Model and Dual-task Learning
- 构建440万条化学反应与逆合成指令数据集
- 联合优化反应生成与逆合成预测,性能领先
- 适合药物研发人员快速生成高活性化合物
化学反应与逆合成预测是药物发现的核心任务。尽管大语言模型在多个领域展现潜力,但直接应用于这些任务仍面临两大挑战:(i) 缺乏大规模化学合成相关指令数据集;(ii) 现有微调策略忽略了反应与逆合成预测之间的紧密关联。为解决上述问题,我们提出ChemDual,一种用于精准化学合成的新颖大语言模型框架。具体而言,鉴于反应与逆合成数据获取成本高昂,ChemDual将分子的反应与逆合成视为相关联的重组与断裂过程,并构建了规模达440万条的指令数据集。此外,ChemDual引入增强版LLaMA,配备多尺度分词器和双任务学习策略,共同优化重组与断裂过程以及反应与逆合成预测任务间的协同。在Mol-Instruction与USPTO-50K数据集上的大量实验表明,ChemDual在反应与逆合成预测上均达到当前最优性能,显著优于传统单任务方法及通用开源大模型。通过分子对接分析,ChemDual生成的化合物表现出多样且强效的蛋白质结合能力,进一步凸显其在药物设计中的巨大潜力。
原文摘要 · Abstract (English)
Chemical reaction and retrosynthesis prediction are fundamental tasks in drug discovery. Recently, large language models (LLMs) have shown potential in many domains. However, directly applying LLMs to these tasks faces two major challenges: (i) lacking a large-scale chemical synthesis-related instruction dataset; (ii) ignoring the close correlation between reaction and retrosynthesis prediction for the existing fine-tuning strategies. To address these challenges, we propose ChemDual, a novel LLM framework for accurate chemical synthesis. Specifically, considering the high cost of data acquisition for reaction and retrosynthesis, ChemDual regards the reaction-and-retrosynthesis of molecules as a related recombination-and-fragmentation process and constructs a large-scale of 4.4 million instruction dataset. Furthermore, ChemDual introduces an enhanced LLaMA, equipped with a multi-scale tokenizer and dual-task learning strategy, to jointly optimize the process of recombination and fragmentation as well as the tasks between reaction and retrosynthesis prediction. Extensive experiments on Mol-Instruction and USPTO-50K datasets demonstrate that ChemDual achieves state-of-the-art performance in both predictions of reaction and retrosynthesis, outperforming the existing conventional single-task approaches and the general open-source LLMs. Through molecular docking analysis, ChemDual generates compounds with diverse and strong protein binding affinity, further highlighting its strong potential in drug design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。