用多任务学习让大模型高效掌握分子理解与生成,实现顶尖逆合成规划。
BioMedGPT-Mol: Multi-task Learning for Molecular Understanding and Generation
- 构建统一指令数据集,通过多任务框架微调通用语言模型。
- 在多个基准上表现卓越,逆合成规划达当前最优水平。
- 适合药物研发、化学生物学领域研究人员使用。
分子在生物医学研究和发现中至关重要,尤其在小分子药物开发中。随着大语言模型的快速发展,尤其是推理模型的出现,探索如何将通用语言模型高效适配于分子科学成为自然方向。本文提出BioMedGPT-Mol,一个支持分子理解与生成任务的分子语言模型。通过整合现有公开指令数据集,构建大规模、综合性且高质量的训练数据集,并采用精心设计的多任务学习框架进行微调。在由LlaSMol、TOMG-Bench和MuMOInstruct合并而成的综合基准上,BioMedGPT-Mol表现出色。实验表明,通过结构化的多任务课程,通用推理模型可被有效且高效地后训练为专业分子语言模型。基于此能力,我们进一步将其应用于多步逆合成规划,在RetroBench上达到领先性能,验证了其作为端到端逆合成规划器的优越性。我们预期该方法可扩展至其他生物医学科学领域。
原文摘要 · Abstract (English)
Molecules play a crucial role in biomedical research and discovery, particularly in the field of small molecule drug development. Given the rapid advancements in large language models, especially the recent emergence of reasoning models, it is natural to explore how a general-purpose language model can be efficiently adapted for molecular science applications. In this work, we introduce BioMedGPT-Mol, a molecular language model designed to support molecular understanding and generation tasks. By curating and unifying existing public instruction datasets, we have assembled a large-scale, comprehensive, and high-quality training dataset. The model is then fine-tuned through a meticulously designed multi-task learning framework. On a consolidated benchmark derived from LlaSMol, TOMG-Bench, and MuMOInstruct, BioMedGPT-Mol achieves remarkable performance. Our experimental results demonstrate that a general-purpose reasoning model can be effectively and efficiently post-trained into a professional molecular language model through a well-structured multi-task curriculum. Leveraging these capabilities, we further apply the model to multi-step retrosynthetic planning, achieving state-of-the-art performance on RetroBench and demonstrating its superior efficacy as an end-to-end retrosynthetic planner. We anticipate that our approach can be extended to other biomedical scientific domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。