arXiv:2502.01074cs.LG2025-02NeurIPS被引 3

构建首个支持任意分子任务的通用模型,统一生成与理解。

Omni-Mol: Multitask Molecular Model for Any-to-any Modalities

  • 按任务类型分类并整合16类分子任务,构建超大规模数据集
  • 提出MoGE架构动态适配任务内在维度,提升多任务学习稳定性
  • 在13项任务上达领先水平,适合分子智能研发人员使用

在分子领域,尽管已有研究探索利用多模态大语言模型构建通用、多任务分子模型,但仍未实现真正意义上的通用性。本文识别出三大挑战:(1) 现有分子任务数据集规模小且覆盖不全;(2) 不同子领域的任务因分布差异大、任务间竞争强,导致联合学习不稳定;(3) 任务间及任务内分子表示对语言空间维度需求不同,难以平衡冗余与不足。为此,本文将现有小分子任务分为四类:Mol2Mol、Mol2Text、Mol2Num和Text2Mol,收集涵盖超过16个任务、140万样本的数据集,为迄今最大分子指令微调数据集。基于大语言模型在化学文献上的充分预训练,提出Omni-Mol框架,统一支持所有小分子任务,兼具生成与理解能力。其核心为提出的MoGE(混合专家)机制,可动态适应不同任务的内在秩。实验表明,该模型在16个任务上实现统一指令微调,并在13个任务上达到当前最优性能,验证了其可扩展性与多功能性。

原文摘要 · Abstract (English)

In the molecular domain, numerous studies have explored the use of multimodal large language models (LLMs) to construct a general-purpose, multi-task molecular model. However, these efforts are still far from achieving a truly universal molecular model. We identify three key challenges in this endeavor: (1) Existing molecular task datasets are typically small in scale and lack comprehensive domain coverage. (2) Tasks from different molecular subfields are difficult to effectively learn jointly through LLMs due to significant distributional shifts and competition among tasks, which introduces instability in the learning process. (3) Both inter-task and intra-task molecular representations demand different intrinsic dimensions in the language space, making it challenging to balance between redundancy and insufficiency in language model representations. To address these challenges, we innovatively categorize existing small-molecule tasks into four types: Mol2Mol, Mol2Text, Mol2Num, and Text2Mol. We then collect a dataset encompassing over 16 tasks with more than 1.4 million samples, making it the largest molecular instruction-tuning dataset to date. Leveraging the extensive pretraining of LLMs on existing chemical literature, we propose a novel multimodal LLM framework, named Omni-Mol, which unifies all small-molecule tasks and supports both molecular generation and understanding. The core of Omni-Mol is our proposed MoGE, which dynamically adapts to the intrinsic rank of different tasks. This mixture-of-experts architecture enhances the model's ability to handle diverse tasks and modalities effectively. Our model achieves unified instruction tuning across 16 tasks and attains state-of-the-art performance on 13 of them. Extensive experiments further demonstrate the scalability and versatility of Omni-Mol.

分子建模多任务学习大模型MoGE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。