arXiv:2505.16326cs.LG2025-05被引 16

ChemMLLM统一建模分子文本、SMILES和图像,提升化学多模态理解与生成能力。

ChemMLLM: Chemical Multimodal Large Language Model

  • 融合文本、SMILES和分子图像的统一多模态架构
  • 在分子图像优化任务中性能超GPT-4o 116.75%
  • 适合药物研发与化学智能生成研究者使用

近年来,多模态大语言模型在多个领域取得显著进展。然而,能够处理跨模态理解与生成的化学多模态模型仍处于探索阶段。为填补这一空白,我们提出ChemMLLM,一个用于分子理解与生成的统一化学多模态大模型。我们设计了涵盖文本、分子SMILES字符串和图像的五项多模态任务,并构建了相应数据集。我们在这些任务上对ChemMLLM与多种主流通用多模态模型及化学大语言模型进行了基准测试。实验结果表明,ChemMLLM在所有评估任务中均表现优异。例如,在分子图像优化任务中,其性能超越最佳基线模型GPT-4o 116.75%(属性提升从1.97增至4.27)。代码已公开于https://github.com/bbsbz/ChemMLLM.git。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have made impressive progress in many applications in recent years. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill this gap, we propose ChemMLLM, a unified chemical multimodal large language model for molecule understanding and generation. Also, we design five multimodal tasks across text, molecular SMILES strings, and image, and curate the datasets. We benchmark ChemMLLM against a range of general leading MLLMs and Chemical LLMs on these tasks. Experimental results show that ChemMLLM achieves superior performance across all evaluated tasks. For example, in molecule image optimization task, ChemMLLM outperforms the best baseline (GPT-4o) by 116.75\% (4.27 vs 1.97 property improvement). The code is publicly available at https://github.com/bbsbz/ChemMLLM.git.

化学大模型多模态分子生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。